Recent testing conducted by blockchain security firm OpenZeppelin revealed that OpenAI's ChatGPT-4, a generative artificial intelligence (AI) model, is not currently capable of effectively auditing smart contracts as human auditors do. The AI model was pitted against OpenZeppelin's Ethernaut security challenge, a wargame consisting of 28 smart contracts to be hacked. While ChatGPT-4 managed to pass a majority of the levels, it faced difficulties with newer levels introduced after its September 2021 training data cutoff date. The absence of a plugin enabling web connectivity during the test hindered its performance.
OpenZeppelin's AI team found that ChatGPT-4 successfully identified vulnerabilities and completed 20 out of the 28 Ethernaut levels. However, it required additional prompts and assistance to solve some levels beyond the initial vulnerability query. Despite these results, the conclusion drawn by Mariko Wakabayashi and Felix Wegener of OpenZeppelin was that ChatGPT-4 cannot currently replace human auditors. They acknowledged the AI model's potential as a tool to enhance the efficiency of smart contract auditors and aid in the detection of security vulnerabilities.
Wakabayashi and Wegener emphasized that smart contract security auditing demands a high level of precision, which existing large language models (LLMs) like ChatGPT are not optimized for. While LLMs excel in generating text and engaging in human-like conversations, the specific requirements of smart contract auditing necessitate tailored training data and output goals for more reliable solutions. While AI can boost auditors' efficiency, OpenZeppelin anticipates continued growth in the number of human auditors employed in the Web3 space due to the ongoing demand for high-quality audits.
