Style-Aware Hierarchical Framework for Museum Artwork Recognition

Main Article Content

Min Kim
Youngbok Cho

Abstract

Automatic museum artwork identification faces significant challenges arising from extremely limited training samples per title—typically a single exemplar image per class—compounded by museum-specific imaging conditions and high intra-class visual variability. This paper presents a style-aware hierarchical recognition framework that addresses these challenges through three coordinated components. First, a domain-specific augmentation pipeline generates ten synthetic variants per original image, simulating spot lighting, lens distortion, motion blur, and partial occlusion. Second, the proposed Style-Aware Feature Fusion (SAFF) module is mounted atop an InceptionResNetV2 backbone; it extracts a 256-dimensional global style embedding via average pooling, generates channel-wise attention weights through a sigmoid-activated dense layer, and reweights the multi-scale feature map in a residual manner to emphasize stylistically discriminative cues such as brushstroke texture, color layering, and compositional rhythm. Third, a two-stage hierarchical classifier first assigns each artwork to one of four departments and then applies a department-specific title-level classifier, reducing the effective class space from 261 to between 42 and 86 classes per model. Evaluated on the Metropolitan Museum of Art benchmark, the framework achieves 86% accuracy at the department level and 84% weighted-average Top-1 accuracy across 261 title classes. The end-to-end hierarchical pipeline yields a final Top-1 accuracy of 75%, substantially outperforming flat baseline models including InceptionResNetV2 (56%), InceptionV3 (45%), and VGG16 (37%). These results demonstrate that combining domain-aware augmentation, style-sensitive feature fusion, and hierarchical decomposition provides a robust and scalable foundation for assistive museum applications.

Article Details

Kim, M., & Cho, Y. (2026). Style-Aware Hierarchical Framework for Museum Artwork Recognition. Journal of Artificial Intelligence Research and Innovation, 66–74. https://doi.org/10.29328/journal.jairi.1001020
Research Articles

Copyright (c) 2026 Kim M, et al.

Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.

Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. In: Pereira F, Burges CJ, Bottou L, Weinberger KQ, editors. Advances in Neural Information Processing Systems (NIPS 2012); 2012 Dec 3–8; Lake Tahoe, NV, USA. Red Hook, NY, USA: Curran Associates; 2012. p. 1097–105.

He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2016); 2016 Jun 27–30; Las Vegas, NV, USA. Piscataway, NJ, USA: IEEE; 2016. p. 770–8. Available from: https://doi.org/10.1109/CVPR.2016.90 DOI: https://doi.org/10.1109/CVPR.2016.90

Senjam SS. Smartphones as assistive technology for visual impairment. Eye. 2021;35(8):2078–80. Available from: https://dx.doi.org/10.1038/s41433-021-01499-w. DOI: https://doi.org/10.1038/s41433-021-01499-w

Sharma P, Kaur R. A survey on smartphone-based assistive technologies for visually impaired people. AIP Conf Proc. 2023;2705(1):040002. Available from: https://dx.doi.org/10.1063/5.0133329. DOI: https://doi.org/10.1063/5.0133329

Lee A. From Braille to Be My Eyes—there's a revolution happening in tech for the blind [Internet]. London, UK: The Guardian; 2017 Jun 26 [cited 2025 Jan 1]. Available from: https://www.theguardian.com.

Sparks H. Houston Museum of Natural Science introduces app for people with low vision [Internet]. New York, NY, USA: Axios; 2025 Feb 11 [cited 2025 Mar 1]. Available from: https://www.axios.com.

Johnson K. AI could change how blind people see the world [Internet]. New York, NY, USA: WIRED; 2023 Jul 5 [cited 2025 Jan 1]. Available from: https://www.wired.com.

Balakrishnan N, Lam SH, Yang P. Using convolutional neural networks to classify and understand artists from the Rijksmuseum. Stanford, CA, USA: Stanford University; 2017. (CS231n Course Report).

Artwork classification and recognition system based on a convolutional neural network. [Research Report]; 2022.

Using convolutional neural networks to classify art genre. [Honours Thesis]. Townsville, QLD, Australia: James Cook University; 2022.

Wang Z, Song H. A fusion model for artwork identification based on convolutional neural networks and transformers. arXiv:2502.18083 [Preprint]. Available from: https://arxiv.org/abs/2502.18083.

Buslaev A, Iglovikov VI, Khvedchenya E, Parinov A, Druzhinin M, Kalinin AA. Albumentations: fast and flexible image augmentations. Information. 2020;11(2):125. Available from: https://dx.doi.org/10.3390/info11020125. DOI: https://doi.org/10.3390/info11020125

Szegedy C, Ioffe S, Vanhoucke V, Alemi AA. Inception-v4, Inception-ResNet, and the impact of residual connections on learning. In: Proceedings of the 31st AAAI Conference on Artificial Intelligence (AAAI 2017); 2017 Feb 4–9; San Francisco, CA, USA. Palo Alto, CA, USA: AAAI Press; 2017. p. 4278–84. DOI: https://doi.org/10.1609/aaai.v31i1.11231

Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, et al. Going deeper with convolutions. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2015); 2015 Jun 7–12; Boston, MA, USA. Piscataway, NJ, USA: IEEE; 2015. p. 1–9. DOI: https://doi.org/10.1109/CVPR.2015.7298594

Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, et al. ImageNet large-scale visual recognition challenge. Int J Comput Vis. 2015;115(3):211–52. Available from: https://dx.doi.org/10.1007/s11263-015-0816-y. DOI: https://doi.org/10.1007/s11263-015-0816-y