This study assessed the accuracy, relevance, and readability of large language models (LLMs) in answering frequently asked questions (FAQs) on corneal transplantation. ChatGPT-3.5, ChatGPT 4o, Gemini, Copilot, and DeepSeek were evaluated using responses to 10 FAQs in Turkish and English. Accuracy was independently rated by two ophthalmologists, while readability was measured using the Ateşman and Çetinkaya indices for Turkish and six established English indices.
Response length was analyzed by character, word, and sentence counts. The proportions of “correct and complete” responses were 72% for ChatGPT-3.5, 80% for ChatGPT 4o, 76% for Gemini, 70% for Copilot, and 71% for DeepSeek (p=0.47). In Turkish, Copilot produced the most complex texts, whereas DeepSeek achieved the highest readability based on the Ateşman index (58.2); according to the Çetinkaya index, ChatGPT 4o had the highest Turkish readability score (19.98).
In English, DeepSeek demonstrated the highest readability across indices, while Gemini consistently generated the most complex outputs. Response length differed significantly between Turkish and English (p<0.001). Overall, LLMs provided largely accurate responses to patient-oriented corneal transplantation queries; however, variability in readability across languages indicates limitations in cross-linguistic communication.
Enhancing linguistic adaptability and content accuracy remains essential for multilingual medical applications.