← Back to repository
2024

Artificial Intelligence (Ai) In Professional Accounting Examination

Osei Adjaye-Gyamfi: FCA, Samuel Koranteng Fianko: P.h.D, Frederick Agropah: CA

1 recorded download0 citations
Download PDF

Abstract

The emergence of artificial intelligence in professional domains has raised critical questions about its impact on accounting certification standards. This March 2024 experimental study compares the performance of 5,698 human candidates against four leading AI models (ChatGPT-3.5, ChatGPT-4, Claude 3, and Gemini) across eight professional accounting subjects in the Ghana Professional Accountancy Examination, providing unprecedented insights into AI’s capabilities in professional certification. Level & Subject Claude GPT-4 GPT-3.5 Gemini Human Best Foundation Level Management Accounting (IMA) 88% 85% 82% 80% 76% Business Management & Information System (BMIS) 94% 92% 100% 85% 83% Business Law (BL) 74% 60% 55% 52% 67% Intermediate Level Audit and Assurance (AA) 88% 99% 82% 83% 87% Public Sector Accounting & Finance (PSAF) 85% 82% 75% 72% 78% Advanced Level Strategic Case Study (SCS) 98% 88% 65% 60% 79% Corporate Reporting (CR) 61% 55% 27% 41% 72% Comparison of the Overall Performance of AI’s Subject-wise Analysis 79.75% 77% 54.38% 50.25% CLAUDE CHATGPT-4 CHATGPT-3.5 GEMINI Impact of Domain-Specific AI Training on AI Performance Overall Improvement +8.66% Average performance increase Strengths of the AIs Advanced AI (Claude and ChatGPT-4) consistently outperformed human benchmarks 100% score in Business Management (ChatGPT-3.5) 99% score in Audit & Assurance (ChatGPT-4) Training improved the performance significantly Key Conclusions Weaknesses of the AIs Basic AI (ChatGPT-3.5 and Gemini) struggled with Corporate Reporting (27-41%) Performance gap in Advanced Level subjects Complex financial analysis challenges Professional judgment limitations Advanced AI models demonstrate strong capabilities in professional accounting examinations, outperforming human benchmarks in structured tasks. However, performance varies significantly between advanced and basic AI models, particularly in complex subjects requiring professional judgment. ivDear Esteemed Members and Stakeholders,