
Artificial Intelligence In Professional Accounting Examination Assessment 2025
1 recorded download0 citations
Abstract
EXECUTIVE SUMMARY
Study Overview
This comprehensive study investigated the effectiveness of four
leading AI models (GPT-4, Claude 3.5, Perplexity, and DeepSeek)
in assessing professional accounting examination scripts compared
to human markers, both with and without standardized marking schemes
across nine accounting subjects.
Research Questions
- What is the effectiveness of AI models in grading Ghana's
professional accounting examination responses? - How do AI models' grading compare with human examiners'
assessments with and without marking schemes? - What are the implications of implementing AI-assisted grading
for professional accounting examinations in Ghana?
Key Findings
AI vs Human Performance Without Marking Schemes
- Most AI models scored more generously than humans.
- Claude showed strongest consistency with humans.
- GPT-4 and Perplexity typically overscored.
- DeepSeek showed most inconsistent patterns, with dramatic
over/underscoring in some areas.
Impact of Marking Schemes on AI-Human Alignment
- Claude demonstrated significant improvement in alignment
with human benchmarks (p=0.014). - GPT-4 generally maintained or increased its deviation
from human standards. - Perplexity and DeepSeek showed inconsistent alignment
patterns with marking schemes.
AI Model Performance Comparison
| AI Model | Alignment with Human Markers | Overall Rating |
|---|---|---|
| Claude 3.5 | Excellent alignment, especially with marking schemes (p=0.014) |
Excellent (90%) |
| Other AI Models | Variable alignment, often overscoring and inconsistent across subjects |
Fair (45%) |