AI-assisted versus traditional systematic review approaches: Evidence from accounting education literature
This study evaluated whether Elicit Pro in Review mode can replicate the coverage achieved through traditional expert-driven systematic literature searches and title/abstract screening stages within the context of accounting education research. The study used a well-established accounting education literature review series as the benchmark study. Elicit’s performance was assessed on repeatability, accuracy, sensitivity and precision metrics. Findings showed that Elicit demonstrates notable strengths in rapid scoping and conceptual exploration but is substantially weaker in all tested performance measurements. Notably, there was minimal overlap between Elicit’s coverage compared to the benchmark study. Moreover, the precision and sensitivity measurements showed that Elicit was consistently in the 1-15% range compared to high 98-100% accuracy measurements for the benchmark study. This study also included a robustness check of Elicit using the GALILEO database. The findings from the robustness check using the GALILEO database reinforced the observed weaknesses of Elicit. Overall, results indicated that Elicit enhances early-stage exploration and thematic discovery but does not replace expert-curated review processes for comprehensive, authoritative literature classification. Instead, the two approaches serve complementary roles within contemporary accounting education research synthesis. Recommendations are provided for future research.
