|
Abstract
|
Dental panoramic X-ray imaging plays a crucial role in diagnosing oral diseases and planning clinical treatments, where accurate tooth segmentation serves as a fundamental preprocessing step for automated dental analysis. In this study, we present a comprehensive analysis of fourteen U-Net-based architectures for semantic tooth segmentation using the recently released STS-Tooth A-PXI dataset, which contains 850 labeled panoramic X-ray images of adult patients. We implemented and evaluated both 2-convolution and 3-convolution per block variants across the U-Net family, including Vanilla U-Net, Dense U-Net, Attention U-Net, SE U-Net, Residual U-Net, R2U-Net, Inception U-Net, Dilated U-Net, UNet++, TransUNet, Swin-
UNet, Efficient U-Net, Double U-Net, and UNet3+. All models were trained using identical preprocessing, augmentation, and evaluation protocols to ensure a fair comparison. Metrics—
accuracy, Dice coefficient, intersection over union (IoU), and mean epoch time—were computed for each architecture. Experimental results demonstrated that U-Net3+ (3conv) achieved the highest test performance with a Dice score of 93.38% and IoU of 87.59%, outperforming other variants in segmentation accuracy and consistency. In contrast, lightweight models such as Vanilla U-Net and Efficient U-Net showed competitive accuracy with significantly reduced computational costs, making them suitable for real-time clinical applications.
|