Acting Like or Thinking Like Human: Doctors and Deep Learning Algorithms Behavioral Output Comparison

Authors

  • Yusuf Altuntaş Department of Orthopedics and Traumatology Service, Şişli Hamidiye Etfal Research and Training Hospital, İstanbul, Türkiye
  • Muhammed Taha Zeren Merkezi Kayıt Kuruluşu A.Ş, Central Securities Depository and Trade Repository of Türkiye, İstanbul, Türkiye
  • Şebnem Özdemir Management Information System Department, İstinye University, İstanbul, Türkiye

Keywords:

Computer vision, deep learning, machine learning, artificial intelligence, artificial neural networks, YOLO, Faster R-CNN, SSD, ANOVA, Tukey pairwise test

Abstract

This study aims compare the behavioral outputs and response durations’ analysis of two distinct groups of medical professionals – general practitioners and orthopedic specialist doctors – utilizing deep learning and computer vision algorithms. This comparison is facilitated through one-way ANOVA and Tukey pairwise comparison tests. The research focuses on femoral proximal fractures which have the highest disability and mortality rates among the elderly population globally. For this research over 1500 femoral proximal region images sourced from over 500 patients aged between 22 for retraining of deep learning algorithms. Updated SSD – Mobilenet v2, Updated Faster R-CNN – Inception v2 and Updated YOLO Darknet v4 algorithms were retrained specifically for femoral proximal fracture detection. In this study, a total of 50 x-ray images comprising femoral proximal fractures and also a total of 50 x-ray images comprising without fractures gathered, in total evaluation dataset established contains 55 x-ray images, underwent scrutiny via one-way ANOVA tests, assessing both tree retrained computer vision deep learning algorithms, as well as physician groups concerning fracture detection success in fractured and non-fractured regions, and for the entirety of the evaluation dataset. Subsequent Tukey pairwise comparison tests were conducted based on the obtained results. Regarding result production durations, it was observed that the average detection time of expert physicians significantly exceeded that of general practitioners, the Updated YOLO algorithm, Updated Faster R-CNN algorithm, and Updated SSD Algorithm. Conversely, no significant disparity was noted between the average result production times of general practitioners and expert physicians. In Tukey pairwise tests evaluating the accuracy rate of fractured regions, the success rate of expert physicians in fractured region accuracy significantly surpassed that of SSD and general practitioners, while insignificantly differing from Faster R-CNN and YOLO Algorithms. For non-fractured region accuracy, the Tukey pairwise test revealed that the accuracy rate of the YOLO algorithm significantly outstripped that of Faster R-CNN and SSD Algorithms, while insignificantly differing from expert and general practitioners. However, it was noted that the success rate of Faster R-CNN and SSD algorithms significantly differed from expert physicians. As conclusion; individual output behavior differs general practitioners and expert physicians but expert physicians show like AI Algorithms similar behaviors in same outputs.

Downloads

Download data is not yet available.

Published

2026-08-31

How to Cite

Altuntaş, Y., Zeren, M. T., & Özdemir, Şebnem. (2026). Acting Like or Thinking Like Human: Doctors and Deep Learning Algorithms Behavioral Output Comparison. Computing and Informatics, 45(4). Retrieved from http://147.213.75.17/ojs/index.php/cai/article/view/7977