BRAIN. Broad Research in Artificial Intelligence and Neuroscience

Volume: 17 | Issue: 3 | Paper number: 1.

Enhancing Image Inpainting Using Hybrid Architecture Combining CNN-Vision Transformer and GAN

Published September 16, 2026
Cite
Simge Coşkun - Burdur Mehmet Akif Ersoy University Bucak Zeliha Tolunay Applied Technology and Management Vocational School (TR), Ali Hakan Işık - Burdur Mehmet Akif Ersoy University, Faculty of Engineering and Architecture (TR),

Abstract

Image inpainting focuses on restoring missing parts of an image in a way that preserves both visual continuity and semantic consistency with the surrounding regions. In this study, a hybrid reconstruction model integrating convolutional neural networks, a Vision Transformer (ViT), and adversarial learning is presented to improve image completion quality. The convolutional layers are responsible for extracting local structural features, whereas the ViT component captures broader contextual relationships and long-range dependencies within the image. To enhance reconstruction performance, a combined loss structure including reconstruction, perceptual, style, edge, and adversarial losses was employed. The proposed approach was tested under various masking conditions through both quantitative and qualitative analyses. The experimental findings indicate that the model achieves strong reconstruction performance, reaching PSNR, SSIM, and LPIPS values of 36.22, 0.9779, and 0.0260, respectively. Additional ablation experiments demonstrate the importance of each architectural component. In particular, excluding the ViT module led to a noticeable decrease in PSNR performance, highlighting the significance of global contextual modeling. Likewise, removing the perceptual loss negatively affected perceptual similarity by increasing the LPIPS score. Visual evaluations indicated that the adversarial learning component contributed to generating more natural, visually coherent, and perceptually realistic reconstructions. Overall, both numerical results and visual evaluations confirm that the proposed framework provides a balanced solution for preserving structural details while maintaining perceptual realism in image inpainting applications.

Academic discipline and sub-disciplines: Artificial Intelligence; Neuroscience; Education

Full Text:

PDF

DOI: http://dx.doi.org/10.70594/brain/17.3/1

Article Overview Video




(C) 2010-2026 EduSoft