YOLOV8n-based gold drawing lacquer painting theme detection method and system

By improving the YOLOv8n network, using the GLU module, SimSPPF and SCSA modules, and combining data enhancement and dynamic weight optimization, the problems of feature extraction and detection accuracy in gilded lacquer painting theme detection are solved, achieving efficient and accurate detection results.

CN120612463APending Publication Date: 2025-09-09FUJIAN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510497865.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

When detecting gilded lacquer painting themes, existing models have problems such as insufficient extraction of complex features of artistic images, low classification and regression accuracy, overlapping detection frames, false detections, and missed detections, resulting in poor detection results.

Method used

An improved YOLOv8n network is used. The Bottleneck module is replaced by the GLU module. SimSPPF and SCSA modules are introduced. The loss function is optimized by combining data enhancement and dynamic weight coefficients to construct a gilt lacquer painting theme detection model for training and testing.

Benefits of technology

It significantly improves the accuracy of gold-drawing and lacquer painting theme detection, enhances the ability to capture micron-level process details, improves the model's computational efficiency and small target detection capabilities, prevents overfitting, and is suitable for real-time detection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612463A_ABST
    Figure CN120612463A_ABST
Patent Text Reader

Abstract

The invention provides a YOLOV8n-based gold-tracing paint drawing theme detection method and system in the technical field of computer vision, and the method comprises the steps: S1, obtaining a large number of historical gold-tracing paint drawing images to construct a data set, and dividing the data set into a training set, a verification set and a test set; s2, creating a gold drawing painting theme detection model, and setting a loss function; s3, training parameters are set, and the gold drawing painting theme detection model is trained through the training set and the training parameters; s4, verifying and testing the trained gold drawing painting theme detection model through the verification set and the test set in sequence; and S5, deploying the gold drawing and painting theme detection model passing the test, and carrying out gold drawing and painting theme detection through the deployed gold drawing and painting theme detection model. The method has the advantage that the accuracy of gold drawing paint drawing theme detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for detecting gilded lacquer painting themes based on YOLO V8n. Background Art

[0002] In the field of computer vision technology, object detection is a fundamental and important task aimed at identifying and locating objects of interest in images or videos. Before the introduction of the YOLO algorithm, deep learning-based object detection methods mainly focused on two-stage models, such as R-CNN, Fast R-CNN, and Faster R-CNN. The common feature of these methods is that they obtain candidate boxes through self-supervised learning or other methods, then input the candidate boxes into a convolutional neural network to predict the category and adjust the candidate boxes to output the predicted boxes. These methods have high computational complexity, high requirements for training resources, and are difficult to train. To address this series of problems, the YOLO (You Only Look Once) algorithm came into being. With its single-stage detection and high efficiency, it has promoted the development of object detection technology.

[0003] Gilded lacquer painting includes gold lacquer painting and gold lacquer wood carving. Currently, many of the unique techniques and craftsmanship of gold lacquer painting are almost lost, making the research and compilation of these gold lacquer paintings particularly urgent and important. Due to the large amount of data, the compilation of gold lacquer paintings requires the rapid and accurate annotation of their types and themes. In the past, such tasks were mostly performed manually, but manual annotation is not only tedious but also inefficient. With the rapid development of science and technology, the use of computer science and technology for annotation is not only efficient, but also produces more accurate and objective results. However, due to the inherent characteristics of gold lacquer painting images, the following problems still exist when using computer deep learning to identify and detect them: 1. Existing models cannot adequately extract the complex features of artistic images, which ultimately affects the accuracy of classification and regression in the detection process; 2. Existing models lack the ability to understand and integrate the complex features of artistic images, resulting in overlapping detection frames, false detections, and missed detections.

[0004] Therefore, how to provide a gilded lacquer painting theme detection method and system based on YOLOV8n to improve the accuracy of gilded lacquer painting theme detection has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for detecting gilded lacquer painting themes based on YOLOV8n, so as to improve the accuracy of gilded lacquer painting themes detection.

[0006] In a first aspect, the present invention provides a method for detecting gilded lacquer painting themes based on YOLO V8n, comprising the following steps:

[0007] Step S1, obtaining a large number of historical gilt lacquer painting images, preprocessing and annotating each of the historical gilt lacquer painting images to construct a data set, and dividing the data set into a training set, a validation set, and a test set based on a preset ratio;

[0008] Step S2: creating a gilded lacquer painting theme detection model based on the YOLOv8n network, and setting the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function;

[0009] The YOLOv8n network consists of a backbone network, a neck network, and a detection head; the backbone network, neck network, and detection head are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain a fused feature; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused feature;

[0010] Step S3: setting the training parameters of the gilt lacquer painting theme detection model, and training the gilt lacquer painting theme detection model using the training set and the training parameters;

[0011] Step S4: verifying and testing the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence;

[0012] Step S5: deploying the gilding and lacquer painting theme detection model that has passed the test, and performing gilding and lacquer painting theme detection using the deployed gilding and lacquer painting theme detection model.

[0013] Furthermore, the step S1 is specifically as follows:

[0014] Acquire a large number of historical gilt lacquer painting images, and perform preprocessing on each of the historical gilt lacquer painting images, including at least image standardization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment;

[0015] Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images;

[0016] A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.5.

[0017] Furthermore, in step S2, the formula of the loss function is:

[0018] L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ;

[0019] Among them, L total Represents the loss value of the loss function; (L dfl +L ciou 0 represents the loss value of the positioning loss sub-function; L cls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

[0020] Furthermore, the step S3 is specifically as follows:

[0021] The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1;

[0022] The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image stitching, and image blending are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

[0023] Furthermore, the step S4 is specifically as follows:

[0024] The trained gilt lacquer painting theme detection model is verified by calculating the accuracy rate using the validation set, and determining whether the accuracy rate is greater than a preset accuracy rate threshold. If not, the verification fails, and the training set is expanded to continue training. If so, the verification succeeds, and:

[0025] The recall rate, F1 score and average precision are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

[0026] In a second aspect, the present invention provides a gilded lacquer painting theme detection system based on YOLOv8n, comprising the following modules:

[0027] A data set construction module is used to obtain a large number of historical gilt lacquer painting images, pre-process and annotate each of the historical gilt lacquer painting images to construct a data set, and divide the data set into a training set, a validation set, and a test set based on a preset ratio;

[0028] A gilded lacquer painting theme detection model creation module is used to create a gilded lacquer painting theme detection model based on the YOLOv8n network, and set the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function;

[0029] The YOLOv8n network consists of a backbone network, a neck network, and a detection head; the backbone network, neck network, and detection head are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain a fused feature; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused feature;

[0030] A gilt lacquer painting theme detection model training module is used to set the training parameters of the gilt lacquer painting theme detection model and train the gilt lacquer painting theme detection model using the training set and the training parameters;

[0031] A model verification and testing module, configured to verify and test the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence;

[0032] The gilded lacquer painting theme detection module is used to deploy the gilded lacquer painting theme detection model that has passed the test, and perform gilded lacquer painting theme detection through the deployed gilded lacquer painting theme detection model.

[0033] Furthermore, the dataset construction module is specifically used to:

[0034] Acquire a large number of historical gilt lacquer painting images, and perform preprocessing on each of the historical gilt lacquer painting images, including at least image standardization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment;

[0035] Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images;

[0036] A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.5.

[0037] Furthermore, in the gilded lacquer painting theme detection model creation module, the formula of the loss function is:

[0038] L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ;

[0039] Among them, L total Represents the loss value of the loss function; (L dfl +L ciou 0 represents the loss value of the positioning loss sub-function; L cls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

[0040] Furthermore, the gilded lacquer painting theme detection model training module is specifically used to:

[0041] The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1;

[0042] The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image stitching, and image blending are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

[0043] Furthermore, the model verification test module is specifically used to:

[0044] The trained gilt lacquer painting theme detection model is verified by calculating the accuracy rate using the validation set, and determining whether the accuracy rate is greater than a preset accuracy rate threshold. If not, the verification fails, and the training set is expanded to continue training. If so, the verification succeeds, and:

[0045] The recall rate, F1 score and average precision are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

[0046] The advantages of the present invention are:

[0047] 1. By obtaining a large number of historical gilded lacquer painting images, preprocessing and annotating each historical gilded lacquer painting image to construct a data set, and dividing the data set into training set, validation set and test set based on a preset ratio; then creating a gilded lacquer painting theme detection model based on the YOLOv8n network, setting the loss function of the gilded lacquer painting theme detection model based on the classification loss subfunction, positioning loss subfunction and confidence loss subfunction, setting the training parameters of the gilded lacquer painting theme detection model, training the gilded lacquer painting theme detection model with the training set and training parameters, verifying and testing the trained gilded lacquer painting theme detection model with the validation set and test set in turn, finally deploying the gilded lacquer painting theme detection model that passes the test, and performing gilded lacquer painting theme detection with the deployed gilded lacquer painting theme detection model; that is, performing gilded lacquer painting theme detection with the gilded lacquer painting theme detection model created based on the YOLOv8n network, the YOLOv8n network consists of a backbone network, a neck network and a detection head; the backbone network is based on the GLU module and the SimSPPF module The invention relates to a novel method for detecting gilded lacquer paintings by using a block-based approach, which is used to extract multi-scale image features from gilded lacquer painting images; a neck network is constructed based on the SCSA module, which is used to fuse the multi-scale image features to obtain fused features; a detection head is used to output a gilded lacquer painting theme detection result containing the predicted probability of the gilded lacquer painting content category based on the fused features; that is, the YOLOv8n network is improved by replacing the original Bottleneck module with the GLU module, so that each feature unit can generate an exclusive gating weight according to the high-frequency details of its local neighborhood, effectively overcoming the defect that the global statistics in the attention mechanism are not sensitive enough to spatial details, significantly improving the ability to capture micron-level process details in gilded lacquer painting images, and avoiding the problem of semantic information loss in the deep feature extraction process; replacing the original SPPF module with the SimSPPF module to improve the computational efficiency of the model and accelerate network convergence; introducing the SCSA module to improve the small target detection capability, suppress shallow feature noise, enhance the positioning accuracy in the detection task, and ultimately greatly improve the accuracy of gilded lacquer painting theme detection.

[0048] 2. By adopting data augmentation methods such as rotation, cropping, Gaussian blur, and random saturation adjustment, the diversity of the dataset is effectively expanded, the generalization ability is enhanced, and the model's robustness to image deformation and lighting changes is significantly improved, preventing overfitting.

[0049] 3. By setting the data set to be divided into training set, validation set and test set in a ratio of 7:1.5:1.5, the amount of validation and test data is increased compared to the traditional 8:1:1, ensuring more rigorous model evaluation.

[0050] 4. The GLU module is used to enhance feature expression capabilities, and the SimSPPF module is combined to improve the efficiency of multi-scale feature extraction, which is superior to the traditional SPPF structure. The feature fusion mechanism of the SCSA module effectively integrates multi-scale features and enhances the detection capabilities of complex textures and small targets. By directly outputting the predicted probability of the content category of the gold lacquer painting, the structure is simple, in line with the lightweight characteristics of YOLOv8n, and suitable for real-time detection needs.

[0051] 5. By combining classification loss, positioning loss and confidence loss, and balancing the optimization directions of different tasks through dynamic weight coefficients, the overall convergence stability of the model is improved.

[0052] 6. By setting the learning rate to 0.0005, the batch size to 16, the dynamic confidence to 0.45, the weight decay coefficient to 0.0005, and the random dropout rate to 0.1, the small target detection capability is effectively enhanced, overfitting is prevented, and the generalization capability is enhanced.

[0053] 7. Stop training if there is no improvement in the mAP indicator after 50 consecutive epochs. This is an early stopping mechanism to avoid overfitting and ensure optimal model performance.

[0054] 8. During the verification phase, the accuracy threshold is used to judge the reliability of the model. If the accuracy threshold is not met, the training set is expanded to form a data closed-loop optimization. During the testing phase, multiple indicators such as recall rate, F1 score, and mAP are comprehensively evaluated to fully reflect the model performance. The model size is also additionally verified to ensure lightweight adaptation to edge deployment scenarios.

[0055] 9. Through the lightweight architecture of YOLOv8n, combined with model compression technology (such as dynamic weight coefficients), it can control computing resource usage while maintaining high accuracy (mAP), making it easy to deploy on embedded devices or mobile terminals.

[0056] 10. Through data enhancement, model architecture innovation, dynamic optimization of loss function and improvement of training strategy, efficient and accurate detection of gold lacquer painting content is achieved, which is both lightweight and practical. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0058] Figure 1 This is a flow chart of a method for detecting gilded lacquer themes based on YOLOV8n in the present invention.

[0059] Figure 2 It is a structural schematic diagram of a gilded lacquer painting theme detection system based on YOLOV8n of the present invention.

[0060] Figure 3 It is a structural diagram of the gilded lacquer painting theme detection model of the present invention.

[0061] Figure 4 This is a structural diagram of the GLU module.

[0062] Figure 5 It is a structural diagram of the SCSA module.

[0063] Figure 6 This is a comparison diagram between the SimSPPF module and the SPPF module. DETAILED DESCRIPTION

[0064] The technical solution in the embodiments of the present application has the following overall idea: gilded lacquer painting theme detection is performed based on a gilded lacquer painting theme detection model created based on an improved YOLOv8n network, and the original Bottleneck module is replaced by the GLU module, so that each feature unit can generate an exclusive gating weight based on the high-frequency details of its local neighborhood, effectively overcoming the defect of insufficient sensitivity of the global statistics to spatial details in the attention mechanism, significantly improving the ability to capture micron-level process details in gilded lacquer painting images, and avoiding the problem of semantic information loss in the deep feature extraction process; replacing the original SPPF module with the SimSPPF module to improve the computational efficiency of the model and accelerate network convergence; introducing the SCSA module to improve the small target detection capability, suppress shallow feature noise, enhance the positioning accuracy in the detection task, and thereby improve the accuracy of gilded lacquer painting theme detection.

[0065] Please refer to Figures 1 to 6 As shown, a preferred embodiment of the present invention's method for detecting gilded lacquer themes based on YOLOV8n includes the following steps:

[0066] Step S1, obtaining a large number of historical gilt lacquer painting images, preprocessing and annotating each of the historical gilt lacquer painting images to construct a data set, and dividing the data set into a training set, a validation set, and a test set based on a preset ratio;

[0067] Step S2: creating a gilded lacquer painting theme detection model based on the YOLOv8n network, and setting the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function;

[0068] By combining classification loss, positioning loss and confidence loss, and balancing the optimization directions of different tasks through dynamic weight coefficients, the overall convergence stability of the model is improved.

[0069] Through the lightweight architecture of YOLOv8n, combined with model compression technology (such as dynamic weight coefficients), it controls computing resource usage while maintaining high accuracy (mAP), making it easy to deploy on embedded devices or mobile terminals.

[0070] The YOLOv8n network consists of a backbone network (Backbone), a neck network (Neck), and a detection head (Head); the backbone network, neck network, and detection head are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain fused features; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused features;

[0071] By adopting the GLU module to enhance feature expression capabilities and combining it with the SimSPPF module to improve the efficiency of multi-scale feature extraction, it is superior to the traditional SPPF structure; through the feature fusion mechanism of the SCSA module, multi-scale features are effectively integrated to enhance the detection capabilities of complex textures and small targets; by directly outputting the predicted probability of the content category of the gold lacquer painting, the structure is simple, in line with the lightweight characteristics of YOLOv8n, and suitable for real-time detection needs.

[0072] Specifically, the GLU module replaces the Bottleneck module in the traditional C2f module. The traditional C2f module passes the input image through the Conv1 convolutional layer, expanding the number of channels to twice the input. It then splits the input image along the channel dimension into two equal sub-feature maps. One sub-feature map is connected to the feature concatenation part via a residual connection, preserving the original information. The other sub-feature map is passed through several Bottleneck modules to extract high-order features. This sub-feature map is then concatenated with the sub-feature map of the original information and finally compressed through the Conv2 convolutional layer to output a feature map of the target dimension. Because the Bottleneck module only has simple convolution operations and a traditional residual structure, and lacks integrated attention mechanisms such as SE and CBAM, it suffers from insufficient feature extraction capabilities and overly coarse extracted features. Therefore, the GLU module was introduced as a replacement.

[0073] The GLU module achieves feature modulation through the element-by-element product of bilinear projections: the main branch performs linear transformations, and the gating branch generates dynamic weights through an activation function. A minimum 3×3 depthwise convolution operation is introduced before the activation layer of the gating branch to construct a gated channel attention mechanism based on fine-grained neighborhood features. The GLU module enables each feature unit to generate its own gating weight based on the high-frequency details of its local neighborhood, effectively overcoming the lack of sensitivity of global statistics to spatial details in the attention mechanism. This local-global feature synergy significantly improves the YOLOv8n model's ability to capture micron-level craft details in gold lacquer painting images while avoiding the loss of semantic information during deep feature extraction.

[0074] By introducing the Sim SPPF module to replace the original SPPF module, the computational efficiency of the model is improved and the network convergence is accelerated.

[0075] There is almost no structural difference between the Sim SPPF module and the SPPF module. Both of them perform channel compression on the input 1×1 convolution layer, halving the number of input channels, and then generate multi-scale features through three serial maximum pooling operations. These features are fused through channel splicing, and finally convolution compression channels are used. The main differences are: 1. The convolution in the Sim SPPF module uses simplified convolution (Sim Conv), while the convolution in the SPPF module contains more complex parameters, such as groups and bias, which makes the number of parameters of the Sim SPPF module lower than that of the SPPF module; 2. The Sim SPPF module uses the ReLU activation function, while the SPPF module uses the SiLU activation function. Due to the existence of exponentials in the SiLU activation function, the Sim SPPF module is superior to the original SPPF module in terms of computing speed.

[0076] In order to improve the model's ability to understand and express complex features and its ability to deeply integrate information, the SCSA (spatial and channel collaborative attention mechanism) module is introduced into the neck network. This not only improves the small target detection ability of the YOLOv8 network, but also suppresses shallow feature noise and enhances the positioning accuracy in detection tasks.

[0077] The SCSA module consists of SMSA (Shared Multi-Semantic Spatial Attention) and PCSA (Progressive Channel Self-Attention). SMSA integrates multi-semantic information and adopts a progressive compression strategy to inject discriminative spatial prior information into the channel self-attention mechanism of PCSA, effectively guiding the recalibration of channel features. At the same time, the single-head self-attention mechanism based on the channel dimension in PCSA enhances the robust interaction between features and further alleviates the multi-semantic information differences between different sub-features in SMSA.

[0078] The purpose of SMSA is to extract the multi-semantic spatial information of each feature and generate spatial priors; spatial attention mainly focuses on the spatial dimensions of different feature maps (i.e., the height and width of the image). By decomposing the features, the focus areas of different semantic information in the spatial dimensions are extracted. First, the input feature map X is decomposed according to height and width, and then global average pooling is applied to each dimension to create two one-dimensional unidirectional sequence structures. These two unidirectional sequence structures are subjected to deep shared one-dimensional convolution with convolution kernels of 3, 5, 7, and 9, respectively, and feature fusion, group normalization, and Sigmoid function activation operations are performed respectively. Finally, after the dimensionality conversion operation, the output feature map X is multiplied element-by-element with the original input feature map X to obtain the output feature map X. S .

[0079] The purpose of PCSA is to establish the interdependence between channels and learn the correlation between feature channels through the channel-level self-attention mechanism. The output of SMSA is used as the input of PCSA. First, the input feature map X S Perform average pooling with a kernel size of 7×7 to obtain the feature map X P :

[0080]

[0081] in,( H,W ) represents the feature map X S Original width and height;( H',W' ) represents the feature map X S Width and height after average pooling; ( 7,7 ) represents the average pooling convolution kernel size;

[0082] Feature Map X P After group normalization, the data are input into the depthwise separable convolution for point-by-point convolution operation. After the basic dimension transformation, the linear projection of the query, key and value is generated. After passing through the channel-by-channel single-head self-attention module, the average pooling is performed again and the sigmoid function is activated. The expression is as follows:

[0083]

[0084] Among them, F proj Denotes the linear projection that generates queries, keys, and values; DWConv1d denotes depthwise separable convolution; C denotes the number of channels; C→C denotes that the number of channels does not change after DWConv1d; (1,1) denotes the convolution kernel size of DWConv1d; Q denotes the query vector; K denotes the key vector; V denotes the value vector; X attn represents a single-head self-attention matrix; T represents transpose; σ() represents the Sigmoid function; X CRepresents the output of the feature map X after SCSA.

[0085] Finally, the expression of the SCSA module is:

[0086] SCSA(X)=PCSA(SMSA(X)).

[0087] Step S3: setting the training parameters of the gilt lacquer painting theme detection model, and training the gilt lacquer painting theme detection model using the training set and the training parameters;

[0088] Step S4: verifying and testing the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence;

[0089] Step S5: deploying the gilding and lacquer painting theme detection model that has passed the test, and performing gilding and lacquer painting theme detection using the deployed gilding and lacquer painting theme detection model.

[0090] Through data enhancement, model architecture innovation, dynamic optimization of loss functions and improved training strategies, efficient and accurate detection of gold lacquer painting content is achieved, which is both lightweight and practical.

[0091] The step S1 is specifically as follows:

[0092] Acquire a large number of historical gilded lacquer painting images and perform preprocessing on each of the historical gilded lacquer painting images, including at least image normalization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment; through data enhancement, the number of images of each gilded lacquer painting content category is substantially balanced;

[0093] By adopting data augmentation methods such as rotation, cropping, Gaussian blur, and random saturation adjustment, the diversity of the dataset is effectively expanded, the generalization ability is enhanced, the model's robustness to image deformation and lighting changes is significantly improved, and overfitting is prevented.

[0094] Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images;

[0095] A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.5.

[0096] By setting the dataset to be divided into training set, validation set and test set in a ratio of 7:1.5:1.5, the amount of validation and testing data is increased compared to the traditional 8:1:1, ensuring more rigorous model evaluation.

[0097] In step S2, the formula of the loss function is:

[0098] L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ;

[0099] Among them, L total Represents the loss value of the loss function; (L dfl +L ciou ) represents the loss value of the positioning loss sub-function; L cls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

[0100] The classification loss adopts a dynamic weight adjustment mechanism, focusing on dealing with the imbalance problem of positive and negative samples. The formula is:

[0101]

[0102] in, Represents the probability value of the predicted category; N represents the number of positive samples; γ and α t Represents a dynamically adjustable weight used to enhance the gradient contribution of difficult samples.

[0103] The formula of distributed focal loss (DFL) is:

[0104]

[0105] Among them, y i Represents the discrete probability distribution label corresponding to the true coordinate; si represents the predicted discrete distribution probability value, which approximates the true coordinate by weighting the probabilities of the left and right adjacent integer values ​​to improve positioning accuracy.

[0106] The formula for the intersection-over-union loss (CIoU) is:

[0107]

[0108] Among them, ρ 2 Represents the square of the distance between the center point of the predicted box and the real box; c represents the diagonal length of the minimum enclosing rectangle; α represents the trade-off parameter to ensure that the gradient direction is consistent with the geometric constraints; v is used to penalize inconsistent aspect ratios:

[0109] (w gt , h gt (represents the actual width and height; (w pred , h pred ) represents the predicted width and height;

[0110] The formula for confidence loss is:

[0111]

[0112] Where M represents the total number of samples; y' i Indicates the true label, the value is 0 or 1; p i Represents the confidence probability.

[0113] The step S3 is specifically as follows:

[0114] The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1;

[0115] By setting the learning rate to 0.0005, the batch size to 16, the dynamic confidence to 0.45, the weight decay coefficient to 0.0005, and the random dropout rate to 0.1, the small target detection capability is effectively enhanced, overfitting is prevented, and the generalization capability is enhanced.

[0116] The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image mosaic and image mixing are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

[0117] By setting the training to stop after 50 consecutive epochs or when the mAP indicator does not improve, the early stopping mechanism is set to avoid overfitting and ensure the best performance of the model.

[0118] The step S4 is specifically as follows:

[0119] The accuracy (precision, P) of the trained gilt lacquer painting theme detection model is calculated using the validation set to determine whether the accuracy is greater than a preset accuracy threshold. If not, the validation fails and the training set is expanded to continue training. If so, the validation succeeds, and:

[0120] The recall rate (R), F1 score (F1 Score) and average precision (mAP) are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

[0121] During the verification phase, the accuracy threshold is used to judge the reliability of the model. If the accuracy threshold is not met, the training set is expanded to form a data closed-loop optimization. During the testing phase, multiple indicators such as recall rate, F1 score, mAP, etc. are comprehensively evaluated to fully reflect the model performance. The model size is additionally verified to ensure lightweight adaptation to edge deployment scenarios.

[0122] The accuracy rate indicates the proportion of detected targets belonging to the target category, that is, the ratio of the number of correctly detected targets to the total number of detected targets, and the expression is P = TP / (TP + FP).

[0123] The recall rate represents the ratio of the number of targets actually detected to the number of all targets detected by the model. It can measure the comprehensiveness of the model and reflect the model's ability to detect targets. The expression is R = TP / (TP + FN).

[0124] Mean Average Precision (mAP) calculates the model's accuracy across different categories and averages them, combining the concepts of Precision and Recall to provide a comprehensive performance evaluation. In object detection tasks in computer vision, the model needs to identify and locate multiple objects in an image. mAP is an overall evaluation of these object detection results. It is calculated by averaging the average precision (AP) for each category and is expressed as:

[0125]

[0126] The F1 score is the harmonic mean of Precision and Recall. It attempts to find a balance between the two to comprehensively evaluate the performance of the model. The expression is: F1 = 2×P×R / (P+R).

[0127] Among them, TP represents the number of samples predicted correctly by the model during verification; FP represents the number of non-samples predicted by the model during verification; FN represents the number of samples not predicted by the model during verification.

[0128] The weight file size refers to the file space required to store the weights of the neural network model (including parameters such as weights and biases). The parameter count refers to the total number of all trainable parameters in the neural network model, including weights and biases. These parameters are adjusted during the training process through optimization algorithms such as gradient descent to minimize the loss function and improve the performance of the model. Both are used to verify the model size and its compatibility with the hardware.

[0129] In specific implementation, the four evaluation metrics can be used: P, R, mAP@0.5, and mAP@0.5-95. mAP@0.5 represents the mAP calculated when the Intersection over Union (IoU) threshold is 0.5; mAP@0.5-95 represents the average mAP under different IoU thresholds within the range of 0.5 to 0.95 (in steps of 0.05).

[0130] A preferred embodiment of the present invention's gilded lacquer painting theme detection system based on YOLO V8n includes the following modules:

[0131] A data set construction module is used to obtain a large number of historical gilt lacquer painting images, pre-process and annotate each of the historical gilt lacquer painting images to construct a data set, and divide the data set into a training set, a validation set, and a test set based on a preset ratio;

[0132] A gilded lacquer painting theme detection model creation module is used to create a gilded lacquer painting theme detection model based on the YOLOv8n network, and set the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function;

[0133] By combining classification loss, positioning loss and confidence loss, and balancing the optimization directions of different tasks through dynamic weight coefficients, the overall convergence stability of the model is improved.

[0134] Through the lightweight architecture of YOLOv8n, combined with model compression technology (such as dynamic weight coefficients), it controls computing resource usage while maintaining high accuracy (mAP), making it easy to deploy on embedded devices or mobile terminals.

[0135] The YOLOv8n network consists of a backbone network, a neck network, and a detection head; the backbone network (Backbone), neck network (Neck), and detection head (Neck) are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain fused features; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused features;

[0136] By adopting the GLU module to enhance feature expression capabilities and combining it with the SimSPPF module to improve the efficiency of multi-scale feature extraction, it is superior to the traditional SPPF structure; through the feature fusion mechanism of the SCSA module, multi-scale features are effectively integrated to enhance the detection capabilities of complex textures and small targets; by directly outputting the predicted probability of the content category of the gold lacquer painting, the structure is simple, in line with the lightweight characteristics of YOLOv8n, and suitable for real-time detection needs.

[0137] Specifically, the GLU module replaces the Bottleneck module in the traditional C2f module. The traditional C2f module passes the input image through the Conv1 convolutional layer, expanding the number of channels to twice the input. It then splits the input image along the channel dimension into two equal sub-feature maps. One sub-feature map is connected to the feature concatenation part via a residual connection, preserving the original information. The other sub-feature map is passed through several Bottleneck modules to extract high-order features. This sub-feature map is then concatenated with the sub-feature map of the original information and finally compressed through the Conv2 convolutional layer to output a feature map of the target dimension. Because the Bottleneck module only has simple convolution operations and a traditional residual structure, and lacks integrated attention mechanisms such as SE and CBAM, it suffers from insufficient feature extraction capabilities and overly coarse extracted features. Therefore, the GLU module was introduced as a replacement.

[0138] The GLU module achieves feature modulation through the element-by-element product of bilinear projections: the main branch performs linear transformations, and the gating branch generates dynamic weights through an activation function. A minimum 3×3 depthwise convolution operation is introduced before the activation layer of the gating branch to construct a gated channel attention mechanism based on fine-grained neighborhood features. The GLU module enables each feature unit to generate its own gating weight based on the high-frequency details of its local neighborhood, effectively overcoming the lack of sensitivity of global statistics to spatial details in the attention mechanism. This local-global feature synergy significantly improves the YOLOv8n model's ability to capture micron-level craft details in gold lacquer painting images while avoiding the loss of semantic information during deep feature extraction.

[0139] By introducing the Sim SPPF module to replace the original SPPF module, the computational efficiency of the model is improved and the network convergence is accelerated.

[0140] There is almost no structural difference between the Sim SPPF module and the SPPF module. Both of them perform channel compression on the input 1×1 convolution layer, halving the number of input channels, and then generate multi-scale features through three serial maximum pooling operations. These features are fused through channel splicing, and finally convolution compression channels are used. The main differences are: 1. The convolution in the Sim SPPF module uses simplified convolution (Sim Conv), while the convolution in the SPPF module contains more complex parameters, such as groups and bias, which makes the number of parameters of the Sim SPPF module lower than that of the SPPF module; 2. The Sim SPPF module uses the ReLU activation function, while the SPPF module uses the SiLU activation function. Due to the existence of exponentials in the SiLU activation function, the Sim SPPF module is superior to the original SPPF module in terms of computing speed.

[0141] In order to improve the model's ability to understand and express complex features and its ability to deeply integrate information, the SCSA (spatial and channel collaborative attention mechanism) module is introduced into the neck network. This not only improves the small target detection ability of the YOLOv8 network, but also suppresses shallow feature noise and enhances the positioning accuracy in detection tasks.

[0142] The SCSA module consists of SMSA (Shared Multi-Semantic Spatial Attention) and PCSA (Progressive Channel Self-Attention). SMSA integrates multi-semantic information and adopts a progressive compression strategy to inject discriminative spatial prior information into the channel self-attention mechanism of PCSA, effectively guiding the recalibration of channel features. At the same time, the single-head self-attention mechanism based on the channel dimension in PCSA enhances the robust interaction between features and further alleviates the multi-semantic information differences between different sub-features in SMSA.

[0143] The purpose of SMSA is to extract the multi-semantic spatial information of each feature and generate spatial priors; spatial attention mainly focuses on the spatial dimensions of different feature maps (i.e., the height and width of the image). By decomposing the features, the focus areas of different semantic information in the spatial dimensions are extracted. First, the input feature map X is decomposed according to height and width, and then global average pooling is applied to each dimension to create two one-dimensional unidirectional sequence structures. These two unidirectional sequence structures are subjected to deep shared one-dimensional convolution with convolution kernels of 3, 5, 7, and 9, respectively, and feature fusion, group normalization, and Sigmoid function activation operations are performed respectively. Finally, after the dimensionality conversion operation, the output feature map X is multiplied element-by-element with the original input feature map X to obtain the output feature map X. S .

[0144] The purpose of PCSA is to establish the interdependence between channels and learn the correlation between feature channels through the channel-level self-attention mechanism. The output of SMSA is used as the input of PCSA. First, the input feature map X S Perform average pooling with a kernel size of 7×7 to obtain the feature map X P :

[0145]

[0146] in,( H,W ) represents the feature map X S Original width and height;( H',W' ) represents the feature map X S Width and height after average pooling; ( 7,7 ) represents the average pooling convolution kernel size;

[0147] Feature Map X P After group normalization, the data are input into the depthwise separable convolution for point-by-point convolution operation. After the basic dimension transformation, the linear projection of the query, key and value is generated. After passing through the channel-by-channel single-head self-attention module, the average pooling is performed again and the sigmoid function is activated. The expression is as follows:

[0148]

[0149] Among them, F proj Denotes the linear projection that generates queries, keys, and values; DWConv1d denotes depthwise separable convolution; C denotes the number of channels; C→C denotes that the number of channels does not change after DWConv1d; (1,1) denotes the convolution kernel size of DWConv1d; Q denotes the query vector; K denotes the key vector; V denotes the value vector; X attn represents a single-head self-attention matrix; T represents transpose; σ( ) represents the Sigmoid function; X C Represents the output of the feature map X after SCSA.

[0150] Finally, the expression of the SCSA module is:

[0151] SCSA(X)=PCSA(SMSA(X)).

[0152] A gilt lacquer painting theme detection model training module is used to set the training parameters of the gilt lacquer painting theme detection model and train the gilt lacquer painting theme detection model using the training set and the training parameters;

[0153] A model verification and testing module, configured to verify and test the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence;

[0154] The gilded lacquer painting theme detection module is used to deploy the gilded lacquer painting theme detection model that has passed the test, and perform gilded lacquer painting theme detection through the deployed gilded lacquer painting theme detection model.

[0155] Through data enhancement, model architecture innovation, dynamic optimization of loss functions and improved training strategies, efficient and accurate detection of gold lacquer painting content is achieved, which is both lightweight and practical.

[0156] The dataset construction module is specifically used for:

[0157] Acquire a large number of historical gilded lacquer painting images and perform preprocessing on each of the historical gilded lacquer painting images, including at least image normalization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment; through data enhancement, the number of images of each gilded lacquer painting content category is substantially balanced;

[0158] By adopting data augmentation methods such as rotation, cropping, Gaussian blur, and random saturation adjustment, the diversity of the dataset is effectively expanded, the generalization ability is enhanced, the model's robustness to image deformation and lighting changes is significantly improved, and overfitting is prevented.

[0159] Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images;

[0160] A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.5.

[0161] By setting the dataset to be divided into training set, validation set and test set in a ratio of 7:1.5:1.5, the amount of validation and testing data is increased compared to the traditional 8:1:1, ensuring more rigorous model evaluation.

[0162] In the gilded lacquer painting theme detection model creation module, the formula of the loss function is:

[0163] L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ;

[0164] Among them, L total Represents the loss value of the loss function; (L dfl +L ciou ) represents the loss value of the positioning loss sub-function; Lcls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

[0165] The classification loss adopts a dynamic weight adjustment mechanism, focusing on dealing with the imbalance problem of positive and negative samples. The formula is:

[0166]

[0167] in, Represents the probability value of the predicted category; N represents the number of positive samples; γ and α t Represents a dynamically adjustable weight used to enhance the gradient contribution of difficult samples.

[0168] The formula of distributed focal loss (DFL) is:

[0169]

[0170] Among them, y i Represents the discrete probability distribution label corresponding to the true coordinate; si represents the predicted discrete distribution probability value, which approximates the true coordinate by weighting the probabilities of the left and right adjacent integer values ​​to improve positioning accuracy.

[0171] The formula for the intersection-over-union loss (CIoU) is:

[0172]

[0173] Among them, ρ 2 Represents the square of the distance between the center point of the predicted box and the real box; c represents the diagonal length of the minimum enclosing rectangle; α represents the trade-off parameter to ensure that the gradient direction is consistent with the geometric constraints; v is used to penalize inconsistent aspect ratios:

[0174] (w gt , h gt ) represents the actual width and height; (w pred , h pred ) represents the predicted width and height;

[0175] The formula for confidence loss is:

[0176]

[0177] Where M represents the total number of samples; y' i Indicates the true label, the value is 0 or 1; pi Represents the confidence probability.

[0178] The gilded lacquer painting theme detection model training module is specifically used to:

[0179] The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1;

[0180] By setting the learning rate to 0.0005, the batch size to 16, the dynamic confidence to 0.45, the weight decay coefficient to 0.0005, and the random dropout rate to 0.1, the small target detection capability is effectively enhanced, overfitting is prevented, and the generalization capability is enhanced.

[0181] The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image mosaic and image mixing are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

[0182] By setting the training to stop after 50 consecutive epochs or when the mAP indicator does not improve, the early stopping mechanism is set to avoid overfitting and ensure the best performance of the model.

[0183] The model verification test module is specifically used to:

[0184] The accuracy (precision, P) of the trained gilt lacquer painting theme detection model is calculated using the validation set to determine whether the accuracy is greater than a preset accuracy threshold. If not, the validation fails and the training set is expanded to continue training. If so, the validation succeeds, and:

[0185] The recall rate (R), F1 score (F1 Score) and average precision (mAP) are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

[0186] During the verification phase, the accuracy threshold is used to judge the reliability of the model. If the accuracy threshold is not met, the training set is expanded to form a data closed-loop optimization. During the testing phase, multiple indicators such as recall rate, F1 score, mAP, etc. are comprehensively evaluated to fully reflect the model performance. The model size is additionally verified to ensure lightweight adaptation to edge deployment scenarios.

[0187] The accuracy rate indicates the proportion of detected targets belonging to the target category, that is, the ratio of the number of correctly detected targets to the total number of detected targets, and the expression is P = TP / (TP + FP).

[0188] The recall rate represents the ratio of the number of targets actually detected to the number of all targets detected by the model. It can measure the comprehensiveness of the model and reflect the model's ability to detect targets. The expression is R = TP / (TP + FN).

[0189] Mean Average Precision (mAP) calculates the model's accuracy across different categories and averages them, combining the concepts of Precision and Recall to provide a comprehensive performance evaluation. In object detection tasks in computer vision, the model needs to identify and locate multiple objects in an image. mAP is an overall evaluation of these object detection results. It is calculated by averaging the average precision (AP) for each category and is expressed as:

[0190]

[0191] The F1 score is the harmonic mean of Precision and Recall. It attempts to find a balance between the two to comprehensively evaluate the performance of the model. The expression is: F1 = 2×P×R / (P+R).

[0192] Among them, TP represents the number of samples predicted correctly by the model during verification; FP represents the number of non-samples predicted by the model during verification; FN represents the number of samples not predicted by the model during verification.

[0193] The weight file size refers to the file space required to store the weights of the neural network model (including parameters such as weights and biases). The parameter count refers to the total number of all trainable parameters in the neural network model, including weights and biases. These parameters are adjusted during the training process through optimization algorithms such as gradient descent to minimize the loss function and improve the performance of the model. Both are used to verify the model size and its compatibility with the hardware.

[0194] In specific implementation, the four evaluation metrics can be used: P, R, mAP@0.5, and mAP@0.5-95. mAP@0.5 represents the mAP calculated when the Intersection over Union (IoU) threshold is 0.5; mAP@0.5-95 represents the average mAP under different IoU thresholds within the range of 0.5 to 0.95 (in steps of 0.05).

[0195] In summary, the advantages of the present invention are:

[0196] 1. By obtaining a large number of historical gilded lacquer painting images, preprocessing and annotating each historical gilded lacquer painting image to construct a data set, and dividing the data set into training set, validation set and test set based on a preset ratio; then creating a gilded lacquer painting theme detection model based on the YOLOv8n network, setting the loss function of the gilded lacquer painting theme detection model based on the classification loss subfunction, positioning loss subfunction and confidence loss subfunction, setting the training parameters of the gilded lacquer painting theme detection model, training the gilded lacquer painting theme detection model with the training set and training parameters, verifying and testing the trained gilded lacquer painting theme detection model with the validation set and test set in turn, finally deploying the gilded lacquer painting theme detection model that passes the test, and performing gilded lacquer painting theme detection with the deployed gilded lacquer painting theme detection model; that is, performing gilded lacquer painting theme detection with the gilded lacquer painting theme detection model created based on the YOLOv8n network, the YOLOv8n network consists of a backbone network, a neck network and a detection head; the backbone network is based on the GLU module and the SimSPPF module The invention relates to a novel method for detecting gilded lacquer paintings by using a block-based approach, which is used to extract multi-scale image features from gilded lacquer painting images; a neck network is constructed based on the SCSA module, which is used to fuse the multi-scale image features to obtain fused features; a detection head is used to output a gilded lacquer painting theme detection result containing the predicted probability of the gilded lacquer painting content category based on the fused features; that is, the YOLOv8n network is improved by replacing the original Bottleneck module with the GLU module, so that each feature unit can generate an exclusive gating weight according to the high-frequency details of its local neighborhood, effectively overcoming the defect that the global statistics in the attention mechanism are not sensitive enough to spatial details, significantly improving the ability to capture micron-level process details in gilded lacquer painting images, and avoiding the problem of semantic information loss in the deep feature extraction process; replacing the original SPPF module with the SimSPPF module to improve the computational efficiency of the model and accelerate network convergence; introducing the SCSA module to improve the small target detection capability, suppress shallow feature noise, enhance the positioning accuracy in the detection task, and ultimately greatly improve the accuracy of gilded lacquer painting theme detection.

[0197] 2. By adopting data augmentation methods such as rotation, cropping, Gaussian blur, and random saturation adjustment, the diversity of the dataset is effectively expanded, the generalization ability is enhanced, and the model's robustness to image deformation and lighting changes is significantly improved, preventing overfitting.

[0198] 3. By setting the data set to be divided into training set, validation set and test set in a ratio of 7:1.5:1.5, the amount of validation and test data is increased compared to the traditional 8:1:1, ensuring more rigorous model evaluation.

[0199] 4. The GLU module is used to enhance feature expression capabilities, and the SimSPPF module is combined to improve the efficiency of multi-scale feature extraction, which is superior to the traditional SPPF structure. The feature fusion mechanism of the SCSA module effectively integrates multi-scale features and enhances the detection capabilities of complex textures and small targets. By directly outputting the predicted probability of the content category of the gold lacquer painting, the structure is simple, in line with the lightweight characteristics of YOLOv8n, and suitable for real-time detection needs.

[0200] 5. By combining classification loss, positioning loss and confidence loss, and balancing the optimization directions of different tasks through dynamic weight coefficients, the overall convergence stability of the model is improved.

[0201] 6. By setting the learning rate to 0.0005, the batch size to 16, the dynamic confidence to 0.45, the weight decay coefficient to 0.0005, and the random dropout rate to 0.1, the small target detection capability is effectively enhanced, overfitting is prevented, and the generalization capability is enhanced.

[0202] 7. Stop training if there is no improvement in the mAP indicator after 50 consecutive epochs. This is an early stopping mechanism to avoid overfitting and ensure optimal model performance.

[0203] 8. During the verification phase, the accuracy threshold is used to judge the reliability of the model. If the accuracy threshold is not met, the training set is expanded to form a data closed-loop optimization. During the testing phase, multiple indicators such as recall rate, F1 score, and mAP are comprehensively evaluated to fully reflect the model performance. The model size is also additionally verified to ensure lightweight adaptation to edge deployment scenarios.

[0204] 9. Through the lightweight architecture of YOLOv8n, combined with model compression technology (such as dynamic weight coefficients), it can control computing resource usage while maintaining high accuracy (mAP), making it easy to deploy on embedded devices or mobile terminals.

[0205] 10. Through data enhancement, model architecture innovation, dynamic optimization of loss function and improvement of training strategy, efficient and accurate detection of gold lacquer painting content is achieved, which is both lightweight and practical.

[0206] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for detecting gilded lacquer painting themes based on YOLO V8n, characterized by: The steps include: Step S1, obtaining a large number of historical gilt lacquer painting images, preprocessing and annotating each of the historical gilt lacquer painting images to construct a data set, and dividing the data set into a training set, a validation set, and a test set based on a preset ratio; Step S2: creating a gilded lacquer painting theme detection model based on the YOLOv8n network, and setting the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function; The YOLOv8n network consists of a backbone network, a neck network, and a detection head; the backbone network, neck network, and detection head are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain a fused feature; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused feature; Step S3: setting the training parameters of the gilt lacquer painting theme detection model, and training the gilt lacquer painting theme detection model using the training set and the training parameters; Step S4: verifying and testing the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence; Step S5: deploying the gilding and lacquer painting theme detection model that has passed the test, and performing gilding and lacquer painting theme detection using the deployed gilding and lacquer painting theme detection model.

2. The method for detecting gilded lacquer painting themes based on YOLO V8n according to claim 1, wherein: The step S1 is specifically as follows: Acquire a large number of historical gilt lacquer painting images, and perform preprocessing on each of the historical gilt lacquer painting images, including at least image standardization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment; Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images; A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.

5.

3. The method for detecting gilded lacquer painting themes based on YOLO V8n according to claim 1, wherein: In step S2, the formula of the loss function is: L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ; Among them, L total Represents the loss value of the loss function; (L dfl +L ciou ) represents the loss value of the positioning loss sub-function; L cls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

4. The method for detecting gilded lacquer painting themes based on YOLO V8n according to claim 1, wherein: The step S3 is specifically as follows: The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1; The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image stitching, and image blending are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

5. The method for detecting gilded lacquer painting themes based on YOLO V8n according to claim 1, wherein: The step S4 is specifically as follows: The trained gilt lacquer painting theme detection model is verified by calculating the accuracy rate using the validation set, and determining whether the accuracy rate is greater than a preset accuracy rate threshold. If not, the verification fails, and the training set is expanded to continue training. If so, the verification succeeds, and: The recall rate, F1 score and average precision are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

6. A YOLOV8n-based gilded lacquer painting theme detection system, characterized by: Includes the following modules: A data set construction module is used to obtain a large number of historical gilt lacquer painting images, pre-process and annotate each of the historical gilt lacquer painting images to construct a data set, and divide the data set into a training set, a validation set, and a test set based on a preset ratio; A gilded lacquer painting theme detection model creation module is used to create a gilded lacquer painting theme detection model based on the YOLOv8n network, and set the loss function of the gilded lacquer painting theme detection model based on the classification loss sub-function, the positioning loss sub-function, and the confidence loss sub-function; The YOLOv8n network consists of a backbone network, a neck network, and a detection head; the backbone network, neck network, and detection head are connected in sequence; the backbone network is constructed based on the GLU module and the SimSPPF module, and is used to extract multi-scale image features from the gold lacquer painting image; the neck network is constructed based on the SCSA module, and is used to fuse the multi-scale image features to obtain a fused feature; the detection head is used to output a gold lacquer painting theme detection result including a predicted probability of the gold lacquer painting content category based on the fused feature; A gilt lacquer painting theme detection model training module is used to set the training parameters of the gilt lacquer painting theme detection model and train the gilt lacquer painting theme detection model using the training set and the training parameters; A model verification and testing module, configured to verify and test the trained gilt lacquer painting theme detection model using the verification set and the test set in sequence; The gilded lacquer painting theme detection module is used to deploy the gilded lacquer painting theme detection model that has passed the test, and perform gilded lacquer painting theme detection through the deployed gilded lacquer painting theme detection model.

7. The YOLO V8n-based gilded lacquer painting theme detection system according to claim 6, characterized in that: The dataset construction module is specifically used for: Acquire a large number of historical gilt lacquer painting images, and perform preprocessing on each of the historical gilt lacquer painting images, including at least image standardization and data enhancement; the data enhancement includes at least rotation, cropping, Gaussian blurring, and random saturation adjustment; Labeling the pre-processed historical gilded lacquer paintings according to their content categories, wherein the content categories include flowers, birds, fish, insects, folk stories, poems, calligraphy and paintings, scenic spots, and auspicious images; A data set was constructed based on the annotated historical gilt lacquer painting images, and the data set was divided into a training set, a validation set, and a test set based on a ratio of 7:1.5:1.

5.

8. The YOLO V8n-based gilded lacquer painting theme detection system according to claim 6, characterized in that: In the gilded lacquer painting theme detection model creation module, the formula of the loss function is: L total =W reg (L dfl +L ciou )+W cls L cls +W obj L obj ; Among them, L total Represents the loss value of the loss function; (L dfl +L ciou 0 represents the loss value of the positioning loss sub-function; L cls Represents the loss value of the classification loss subfunction; L obj Represents the loss value of the confidence loss sub-function; L dfl represents the distribution focus loss; L ciou W represents the intersection-over-union loss; reg 、W cls and W obj Both represent dynamic weight coefficients.

9. The YOLO V8n-based gilded lacquer painting theme detection system according to claim 6, characterized in that: The gilded lacquer painting theme detection model training module is specifically used to: The gilt lacquer painting theme detection model is set to at least include training parameters of learning rate, batch size, dynamic confidence, weight decay coefficient and random dropout rate, with the learning rate set to 0.0005, the batch size set to 16, the dynamic confidence set to 0.45, the weight decay coefficient set to 0.0005, and the random dropout rate set to 0.1; The gilt lacquer painting theme detection model is trained using the training set and training parameters. During the training process, mixed precision training, four-image stitching, and image blending are combined, and the loss function is continuously optimized until 50 consecutive epochs are trained or the mAP indicator does not improve.

10. The YOLO V8n-based gilded lacquer painting theme detection system according to claim 6, characterized in that: The model verification test module is specifically used to: The trained gilt lacquer painting theme detection model is verified by calculating the accuracy rate using the validation set, and determining whether the accuracy rate is greater than a preset accuracy rate threshold. If not, the verification fails, and the training set is expanded to continue training. If so, the verification succeeds, and: The recall rate, F1 score and average precision are calculated using the test set to test the successfully verified gilded lacquer painting theme detection model. If the test is successful, the model size of the gilded lacquer painting theme detection model is verified.

Citation Information

Cited By

  • Wild animal detection method fusing unmanned aerial vehicle thermal infrared image and visible light image

    CN121305621A

  • Wildlife detection method fusing unmanned aerial vehicle thermal infrared image and visible light image

    CN121305621B