Long-horned beetle image detection and recognition system and method
Through YOLOv5s and the improved ResNet50 model combined with the CBAM attention mechanism, the accuracy and recognition accuracy problems in Tianniu image detection and recognition are solved, and efficient and accurate Tianniu recognition is achieved. It is suitable for real-time applications on mobile terminals and improves Tianniu quarantine work efficiency.
Patent Information
- Application Number
- CN202510313462.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems with low detection accuracy and insufficient accuracy of similar species recognition in complex contexts in the detection and recognition of long-term elbow images, especially when morphological identification relies on low-resolution pictures.
The YOLOv5s model is used for positioning the position of Tianniu, combined with the improved ResNet50 model, integrated into the CBAM attention mechanism, and integrated into the applet through the ONNX Runtime reasoning engine and the FastAPI framework to achieve lightweight deployment and support real-time mobile recognition.
It improves the accuracy and recognition accuracy of Tianniu detection, and achieves efficient identification of 689 types of Tianniu, with an accuracy rate of 97.82%, a recall rate of 97.24%, and an average accuracy of 98.94%. It is suitable for mobile applications, reducing the workload of data labeling and improving work efficiency.
Smart Images

Figure CN120279577A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a longicorn beetle image detection and recognition system and method. Background Art
[0002] With the continuous development of global trade, China's ports are facing increasingly severe challenges of alien biological invasions. As one of the most harmful groups among alien invasive species, longicorn beetles are frequently intercepted by the customs quarantine department from imported timber. From 2003 to 2018, a total of 142,476 species of longicorn beetles were intercepted during the entry quarantine in China, involving 332 genera and 578 species. Among them, as many as 39 species (genera) and 7,552 species are listed in the Catalogue of Alien Plant Quarantine Pests in China. Longicorn beetles can damage living trees, spread dangerous diseases, and can also damage forest products, reducing the technological value of wood. A large number of alien longicorn beetles entering China's ports pose a huge potential threat to China's agricultural and forestry production and ecological environment.
[0003] At present, the quarantine and identification of longicorn beetles still mainly rely on morphological identification, supplemented by molecular biological identification. However, there are more than 38,000 species of longicorn beetles under the superfamily Cerambycidae, and all the pictures in the current domestic and foreign longicorn beetle identification materials are black and white line drawings. Only more than 100 species of adult longicorn beetles intercepted at China's ports are included in the existing materials. However, the pictures of identification characteristics, etc., are mostly low-resolution color photos limited by the shooting technology of that year. These far from meet the morphological identification needs of longicorn beetles intercepted at China's ports. And without the guidance of experienced experts, it is very difficult for novices to get started and accurately identify the adult longicorn beetles encountered in the quarantine work through morphological characteristics.
[0004] With the rise of deep learning technology, models based on convolutional neural networks (CNNs) have shown excellent performance in image recognition tasks. For example, some studies have applied basic CNN models to insect recognition and achieved better results than traditional methods. In the prior art, an improved deep learning object detection model YOLOv4-TIA is used to automatically extract features and perform recognition and detection on forestry insect images, and its mean average precision (mAP) can reach 98.8%. In the prior art, a hierarchical classification model using a deep convolutional network is used to identify and classify 70 typical species frequently intercepted in imported timber at the three classification levels of family, genus, and species. The average recognition accuracies of this model for 70 types of harmful organisms at the family, genus, and species levels are 97.71%, 95.85%, and 86.92% respectively. However, a single CNN model still has limitations when facing the dual tasks of longicorn beetle image detection and recognition, such as low detection accuracy of longicorn beetles in complex backgrounds and room for improvement in the recognition accuracy of similar longicorn beetle species. Summary of the Invention
[0005] To address the problems mentioned in the background art, the present invention proposes a longicorn beetle image detection and recognition system and method. By optimizing the model structure and introducing an attention mechanism, the detection accuracy of longicorn beetles and the fine-grained classification ability in complex environments are improved, and it is integrated into a lightweight mini-program to achieve real-time application in the port quarantine scenario.
[0006] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] A longicorn beetle image detection and recognition system, comprising:
[0008] Data preprocessing module: Collect the required data and label the collected data set.
[0009] Longicorn beetle detection module: Based on the YOLOv5s model, accurately locate the position of the longicorn beetle in the image.
[0010] Data cropping module: Enlarge the detection box according to the detected position of the longicorn beetle and then crop the image to obtain a longicorn beetle target image data set.
[0011] Longicorn beetle recognition module: Based on the improved ResNet50 model, incorporating the CBAM attention mechanism, for high-precision recognition of longicorn beetle species.
[0012] Lightweight deployment: Through the ONNX Runtime inference engine and the FastAPI framework, integrate the model into the mini-program to support real-time recognition on mobile devices.
[0013] Preferably, the specific content of the data preprocessing module is as follows:
[0014] Data collection: Collect specimen images and ecological images including longicorn beetles.
[0015] Data preprocessing: Randomly extract a certain number of images from the collected image library as a sample set, use the LabelMe tool to label the positions of longicorn beetles in the images, and convert the labeled data set into the YOLOv5 format; finally, divide the data set into a training set and a validation set according to a ratio.
[0016] Preferably, the specific content of the longicorn beetle detection module is as follows:
[0017] Based on the YOLOv5 model architecture, use two longicorn beetle detectors. One detector selects the YOLOv5l model structure to label the positions of longicorn beetles in the data set; the other detector needs to select the YOLOv5s model structure to achieve fast detection of longicorn beetles.
[0018] During the training process, a comprehensive evaluation is carried out after each round of training, and the model generated in the current round is saved synchronously.
[0019] Preferably, the specific content of the data cropping module is as follows:
[0020] Data cropping: Use the previous longhorn beetle detector to detect each image in the entire dataset. According to the detected positions, expand the detection boxes into squares and then crop the images to obtain longhorn beetle target images, which are saved in the corresponding species folders; finally, a longhorn beetle target image dataset is formed.
[0021] Data cleaning: Use the Cleanlab library to automatically clean the longhorn beetle target image dataset.
[0022] Preferably, the specific content of the longhorn beetle recognition module is as follows:
[0023] The improved ResNet50 model includes network structure adjustment, attention mechanism fusion, and classification layer optimization. Specifically:
[0024] Network structure adjustment: Change the stride of the last stage in the ResNet50 module to 1 to retain finer feature maps.
[0025] Attention mechanism fusion: Embed the CBAM module after the deformable convolution, and focus on the key texture and color features of longhorn beetles through channel and spatial attention weighting.
[0026] Classification layer optimization: Use the GAP-BN-FC-BN structure to replace the traditional fully connected layer, and combine the label smoothing cross-entropy loss function to alleviate class imbalance.
[0027] Preferably, the CBAM attention mechanism includes channel attention and spatial attention; the feature maps are sequentially processed deeply by channel attention and spatial attention to learn the importance degrees of each feature channel and spatial dimension.
[0028] The calculation process of channel attention is as follows:
[0029] ,
[0030] where M c (F) represents the output of the channel attention module, represents the Sigmoid activation function; MLP represents the multi-layer perceptron, AvgPool(F) represents global average pooling of the input feature map F; MaxPool(F) represents global maximum pooling of the input feature map F; W0 and W1 represent the weight matrices of two convolutional layers in the MLP; represents the result after global average pooling for each channel of the feature map F; represents the result after global maximum pooling for each channel of the feature map F.
[0031] Preferably, the calculation formula for spatial attention is as follows:
[0032] ,
[0033] where M s (F) represents the output of the spatial attention module, σ represents the Sigmoid activation function, and f 7×7 represents a 7×7 convolution operation, AvgPool(F) represents global average pooling on the input feature map F, and MaxPool(F) represents global max pooling on the input feature map F; [AvgPool(F); MaxPool(F)] represents concatenating the results of global average pooling and global max pooling along the channel dimension; and represent the feature maps after global average pooling and global max pooling respectively.
[0034] Preferably, the output result of the CBAM attention mechanism will first go through a global average pooling operation: by taking the average of the entire feature map in the spatial dimension, the feature map is converted into a feature vector with global information. The specific calculation formula is as follows:
[0035] ,
[0036] where represents the output of the global average pooling operation on the feature map output by CBAM, H represents the height of the feature map, W represents the width of the feature map; X represents the input feature map; i, j represent the spatial coordinates of the feature map; n represents the nth sample in the batch; c represents the cth channel of the feature map.
[0037] Preferably, batch normalization is performed on the output of global average pooling. The specific process is as follows:
[0038] Calculate batch statistics, including calculating the mean and calculating the variance;
[0039] Calculate the mean, calculated channel by channel. The specific calculation formula is as follows:
[0040] ,
[0041] where μ c represents the mean of the cth channel in the current batch; N represents the batch size; represents the output data of global average pooling;
[0042] Calculate the variance. The specific calculation formula is as follows:
[0043] ,
[0044] where Represents the variance of the c-th channel in the current batch. Represents the smoothing term; Represents the output data of global average pooling; μ c Represents the mean of the c-th channel in the current batch; N represents the batch size;
[0045] Normalization, specifically:
[0046] ,
[0047] Among them, Represents the output result after normalization; Represents the output result after global average pooling; μ c Represents the mean of the c-th channel in the current batch; Represents the variance of the c-th channel in the current batch;
[0048] Affine transformation, specifically:
[0049] ,
[0050] Among them, Represents the output result after normalization and affine transformation; γ c Represents the scaling factor; β c Represents the offset; Represents the output result after normalization.
[0051] A longhorn beetle image detection and recognition method, according to the longhorn beetle image detection and recognition system described in any of the above claims, includes the following steps:
[0052] S1: Obtain the required image data and annotate it;
[0053] S2: Use the longhorn beetle detector to detect longhorn beetle targets in the input image and perform target cropping;
[0054] Extract images from the collected image data as a sample set, and accurately annotate this sample set so that the drawn bounding box tightly surrounds the boundary of the longhorn beetle target to locate the position of the longhorn beetle;
[0055] Use the longhorn beetle detector to detect the data set, expand the detection box into a square according to the detected position and then crop the image, so as to obtain the longhorn beetle target image and save it to the corresponding species folder;
[0056] Detect the longhorn beetle target area in the image through the YOLOv5s model;
[0057] S3: Use the longhorn beetle recognizer to identify and classify the cropped and cleaned images;
[0058] S31: Automatically clean the longicorn target image dataset to obtain a relatively clean dataset;
[0059] S32: Input the target area into the improved ResNet50 model and combine it with the CBAM attention mechanism to output the species classification results;
[0060] S4: Implement image uploading, real-time reasoning and result display in the mini program.
[0061] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0062] (1) High-precision detection and recognition: The present invention accurately separates the longhorn beetle detection model and the recognition model, and has powerful recognition capabilities, capable of accurately identifying up to 689 species of longhorn beetles. In the test phase, a test set consisting of more than 70,000 longhorn beetle images was used, showing excellent performance, with an accuracy rate (P) of 97.82%, a recall rate (R) of 97.24%, and a mean average precision (mAP) of 98.94%. These data fully demonstrate the efficiency and accuracy of longhorn beetle recognition.
[0063] (2) High efficiency and lightweight: The model of the present invention has the remarkable feature of being lightweight. If the algorithm is directly deployed on a mobile phone terminal (this does not refer to the WeChat applet that relies on background operation, but emphasizes that the model can run independently on the local mobile phone), then further exploring the optimization path of the lightweight model has great potential and value, which can give full play to the local computing resources of the mobile phone and realize efficient longicorn identification applications; the applet uses the ONNX Runtime inference engine and can complete image recognition within 400 milliseconds. The overall response time is within 1 second, which is suitable for mobile applications.
[0064] (3) Dataset optimization: With the help of pre-labeling technology and cleanlab tools, the present invention can effectively reduce the large amount of manpower and time costs consumed in the data labeling process, and significantly improve work efficiency; secondly, it innovatively separates the longicorn beetle identification function from the longicorn beetle detection system, thereby creating conditions for the use of more diverse identification optimization strategies, which effectively promotes the improvement of identification accuracy; reduces the workload of data labeling, and at the same time improves the diversity and quality of the data set.
[0065] (4) Wide application scenarios: The present invention can organically combine fine-grained classification-related algorithms to more accurately classify and identify longicorn beetles, explore more detailed feature differences in longicorn beetle images, further improve the accuracy and professionalism of the model in longicorn beetle identification tasks, and improve the efficiency of longicorn beetle quarantine work. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1It is the flowchart of the longhorn beetle recognition algorithm of the present invention;
[0067] Figure 2 They are partial longhorn beetle pictures for longhorn beetle species recognition of the present invention, where A - E are all Plagionotus arrowianus;
[0068] Figure 3 It is the schematic diagram of the traditional ResNet50 model of the present invention;
[0069] Figure 4 It is the schematic diagram of the improved ResNet50 model of the present invention;
[0070] Figure 5 It is the schematic diagram of the interface for identifying longhorn beetles in the insect sea of the present invention;
[0071] Figure 6 It is the schematic diagram of Interface 1 for AI recognition of the present invention;
[0072] Figure 7 It is the schematic diagram of Interface 2 for AI recognition of the present invention;
[0073] Figure 8 It is the schematic diagram of Interface 3 for AI recognition of the present invention. Specific embodiments
[0074] The following combines specific embodiments to further clarify the present invention. The embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0075] As Figure 1 shown, the longhorn beetle image detection and recognition system and method provided in this embodiment are based on the longhorn beetle detection model of YOLOv5s and the longhorn beetle recognition model based on the improved ResNet50. First, based on the YOLOv5s model, using its efficient and fast advantages in object detection, the parameters are optimized for the longhorn beetle image features to accurately locate the position of the longhorn beetle in the image; at the same time, the improved ResNet50 recognition model integrates the CBAM attention mechanism to effectively mine the key features of the longhorn beetle appearance, overcome the problem of similar species recognition, and achieve high-precision longhorn beetle species recognition and classification. On this basis, the above two models are integrated into a small program. Through a simple interface design, it is convenient for port staff to upload longhorn beetle images conveniently for real-time detection and recognition. Verified by a large number of test samples and applied in the actual port scenario, this small program can accurately and quickly detect and recognize longhorn beetle species in a complex environment, significantly improve the efficiency of port longhorn beetle quarantine work, provide strong technical support for preventing the invasion of foreign longhorn beetles and protecting ecological security, and play a positive role in promoting the development of intelligent identification technology for harmful organisms at ports.
[0076] Deployment of Longhorn Beetle Detection and Recognition System: The YOLOv5s model is used in the longhorn beetle detector, and the improved ResNet50 model with an integrated attention mechanism is adopted in the longhorn beetle recognizer. In the design of the system's backend architecture, ONNX Runtime is selected as the model inference engine to accelerate the running speed of the model and ensure the efficient and stable operation of the system when processing a large amount of longhorn beetle image data.
[0077] In this embodiment, FastAPI is used as the Web framework. During the development of the web front-end, the bootstrap framework is mainly used. With its rich component library and flexible layout design functions, a simple and beautiful front-end interaction page with a good user experience is created, enabling users to conveniently upload longhorn beetle images, view detection results, etc. For the WeChat mini-program part, the native development method is adopted, making full use of the native features and interfaces of the WeChat mini-program platform to provide users with a more convenient and smooth mobile application experience, allowing users to use the longhorn beetle detection and recognition service through the mobile phone WeChat at any time and anywhere, greatly expanding the application scenarios and user group scope of the system.
[0078] A longhorn beetle image detection and recognition system mainly includes:
[0079] Data preprocessing module: Collect the required data and perform data annotation on the collected dataset.
[0080] Data collection: The collection includes specimen images taken, ecological images taken, images from citizen science communities (such as research-level data on iNaturalist (a natural observer platform)), etc. Additionally, to reduce the impact of class imbalance on recognition performance, there should be at least 5 instances for each species participating in the classification training. During the collection process, attention should be paid to the diversity of images, covering aspects such as scene diversity, lighting diversity, scale diversity, and category diversity.
[0081] The longhorn beetle image dataset adopts a specific storage format for easy management and use. Specifically, the images of the same species are uniformly stored in the same folder, based on the fact that it is relatively rare for a single image to contain multiple species under normal circumstances, and such complex situations have been deliberately avoided during the image collection process. For the extremely few images containing multiple species, they are all centrally stored in a specially set specific folder for unified management.
[0082] Figure 1Some representative longicorn beetle images in the dataset are shown. Through these images, the richness and diversity of longicorn beetle images in the dataset can be intuitively felt. This dataset is jointly composed of specimens and ecological images of 689 species of longicorn beetles under Cerambycidae, and the total number of its images exceeds 70,000, providing a solid data foundation for the training of longicorn beetle detection and recognition models.
[0083] Data preprocessing: In the preprocessing stage of the dataset, first, a certain number of images are randomly selected from the collected image library. Here, 10,000 images are selected as the sample set. Subsequently, with the help of the powerful LabelMe tool, the precise location of each longicorn beetle target in these images is marked. During the marking process, specific marking principles are strictly followed, that is, the drawn bounding box must closely surround the boundary of the longicorn beetle object, ensuring that no effective area of the longicorn beetle is missed, thus guaranteeing the integrity and accuracy of the data, and at the same time, avoiding including redundant background areas to reduce unnecessary interference information.
[0084] After completing the marking work, the marked dataset is further converted into a specific marking format adapted to the YOLOv5 model, so that the subsequent model can smoothly read and process the data. Finally, the entire dataset is divided into a training set and a validation set according to a ratio of 9:1. The training set occupies a larger proportion and is mainly used for the training and learning process of the model, enabling the model to fully absorb the feature information and regular patterns in the data; while the validation set is used to conduct phased evaluations and verifications of the model's performance during the model training process, promptly discover problems such as overfitting or underfitting that may exist in the model, and accordingly optimize and adjust the model to ensure that the model has good generalization ability and stability, and can still maintain a high accuracy and reliability when facing unknown data.
[0085] Longicorn beetle detection module: Based on the YOLOv5s model, accurately locate the position of the longicorn beetle in the image, specifically including:
[0086] In the longhorn beetle detection task, the YOLOv5 model architecture carefully built based on the PyTorch framework is adopted. YOLOv5 is initialized with a model trained on the COCO dataset. This process covers the training procedures of two detectors, and each has clear uses and objectives. One detector focuses on the pre-annotation of the positions of longhorn beetles in the subsequent dataset. Given the extremely strict requirements for model accuracy in this application scenario, a relatively large model structure, namely YOLOv5l, is selected to ensure that the position information of longhorn beetles in the dataset can be accurately located, laying a solid foundation for subsequent processing and analysis. The other detector mainly serves the online application scenario. Considering the special requirements of the online application for the model response time, while ensuring a certain detection accuracy, the processing duration must be shortened as much as possible. Therefore, this detector adopts a relatively small and more lightweight model structure, YOLOv5s, to achieve fast and effective longhorn beetle detection. During the training process, a comprehensive evaluation is carried out after each round of training, and the model generated in the current round is saved synchronously. Finally, by carefully comparing and analyzing the performance of each model on the validation set, the model that shows the best comprehensive performance on the validation set is selected and determined as the core model adopted in the actual application of the entire system. In terms of the parameter settings for model training, the model input size is set to 640x640, the batch size for each training is 32, and the total number of epochs for training is set to 300. The reasonable configuration of these parameters aims to balance the training efficiency and performance of the model, ensuring that the model can fully learn the key features and pattern information of longhorn beetle images, so as to exert a stable and excellent detection effect in actual applications.
[0087] Data cropping module: Expand the detection box according to the detected position of the longhorn beetle and then crop the image to obtain a longhorn beetle target image dataset;
[0088] Data cropping: Use the previously obtained larger longhorn beetle detector to detect each image in the entire dataset (except for images containing multiple species). Expand the detection box into a square according to the detected position and then crop the image to obtain the longhorn beetle target image and save it to the corresponding species folder. Finally, a longhorn beetle target image dataset is formed (the storage form is a collection of folders, each folder represents a longhorn beetle species, and all the cropped images of that species are stored under the folder). In this way, only the positions of longhorn beetle targets in a small number of images need to be annotated, and the positions of longhorn beetle targets in the remaining images can be obtained using the pre-trained model, so the annotation workload can be reduced.
[0089] Data cleaning: Since the longhorn beetle detector may have false detection situations, to alleviate this problem, the cleanlab library is used to automatically clean the longhorn beetle target image dataset, and finally a relatively clean dataset is obtained.
[0090] Longhorn beetle recognition module: Based on the improved ResNet50 model with the CBAM attention mechanism integrated, it is used to accurately identify the species of longhorn beetles.
[0091] Improved ResNet50 object recognition model: ResNet50 is used as the basic model for the longhorn beetle recognition task. The advantage of the ResNet50 model is that it solves the problem of gradient disappearance through residual connections, has strong feature learning ability, good transfer learning effect, high model stability, wide application range and fast convergence speed.
[0092] Algorithm details:
[0093] 1) Data augmentation strategy;
[0094] This algorithm adopts a variety of data augmentation methods to expand the diversity of data samples, thereby improving the generalization ability of the model, including: random translation, random horizontal and vertical flipping, random rotation, random color jitter, random zoom blur, random image quality level transformation, random grayscale conversion, ResizeMix, Cutout.
[0095] 2) Network structure design;
[0096] The Backbone of the network selects ResNet50, which has strong feature extraction ability. Initializing with the model parameters pre-trained on the ImageNet dataset can accelerate the training process of the model and improve the initial performance of the model. Based on the structure of ResNet50, the model can effectively extract multi-level features of longhorn beetle images, laying a solid foundation for subsequent classification and recognition tasks.
[0097] 3) Selection of loss function;
[0098] The cross-entropy loss function with label smoothing regularization (LabelSmoothingRegularization, LSR) is adopted. The label smoothing technique avoids the model from overfitting the training data by smoothing the true labels to a certain extent, reducing the risk of overfitting, and enabling the model to have better generalization performance when facing unknown data. The cross-entropy loss is used to measure the difference between the model prediction result and the true label, guiding the model to continuously optimize the parameters during the training process and improve the prediction accuracy. Specifically:
[0099] Label smoothing processing: Let the true label y be a one-hot encoded vector, and label smoothing adjusts each element of y to:
[0100] ,
[0101] Among them, is the smoothing factor, C is the total number of categories, and y i is the i-th element of the true label y.
[0102] Calculate the cross-entropy loss. The specific calculation formula is:
[0103] ,
[0104] where p i is the probability distribution predicted by the model; C is the total number of categories; y i ′ is the i-th element of the label y after smoothing.
[0105] 4) Training plan planning;
[0106] Optimizer selection: Select the Stochastic Gradient Descent (SGD) optimizer, which has good convergence and stability when dealing with large-scale data.
[0107] Learning rate strategy: Adopt the learning rate warm-up mechanism. In the first 5 epochs (referring to the process of completing a full forward and backward propagation on the entire training dataset) at the beginning of training, gradually increase the learning rate from a lower value to the initial set value of 0.01, which helps the model better adapt to parameter updates in the initial stage of training. Subsequently, adopt the cosine annealing learning rate decay strategy. As the number of training epochs increases, the learning rate gradually decays to the final learning rate of 1e-5 according to the law of the cosine function.
[0108] This learning rate adjustment strategy can balance the convergence speed and accuracy of the model during training and avoid the model falling into a local optimal solution in the later stage of training.
[0109] Model input and training parameters: Set the model input size to 256x256, the batch size (batch size) for each input is 128, and the total number of training epochs is 100. By reasonably setting these parameters, while ensuring that the model can fully learn image features, improve the training efficiency, and enable the model to achieve better training results within limited computing resources and time.
[0110] 5) Model optimization;
[0111] The improved ResNet50 model is shown in Figure 2 . At the network structure level, a key adjustment is made to the setting of the last stage. In the traditional ResNet50 model, the stride of this stage is usually 2, while in the improved model, it is set to 1. This change has an important impact on the size of the feature map and subsequent information transmission, enabling the feature information to be processed and transmitted at a finer scale in this stage and reducing the information loss caused by a large stride.
[0112] Furthermore, settings are made in the output processing flow of the improved ResNet50 model at the last stage. Its output results will first go through the global average pooling (GAP) operation. By taking the average of the entire feature map in the spatial dimension, the feature map is effectively converted into a feature vector with global information, greatly compressing the data volume and retaining the key global feature information. The specific calculation formula is:
[0113] ,
[0114] where, represents the output of the global average pooling operation on the feature map output by CBAM, H represents the height of the feature map, W represents the width of the feature map; X represents the input feature map; i, j represent the spatial coordinates of the feature map; n represents the nth sample in the batch; c represents the cth channel of the feature map.
[0115] Then, batch normalization (BN) processing is carried out, which can accelerate the convergence speed of the model and improve the stability of the model. Specifically, it includes:
[0116] 1. Calculate batch statistics, specifically:
[0117] Calculate the mean, channel by channel. The specific calculation formula is:
[0118] ,
[0119] where, μ c represents the mean of the cth channel in the current batch, and the calculation method is to take the average of the values of this channel for all samples in the batch; N represents the batch size; represents the output data of the global average pooling;
[0120] Calculate the variance. The specific calculation formula is:
[0121] ,
[0122] where, represents the variance of the cth channel in the current batch, represents the smoothing term; represents the output data of the global average pooling; μ c represents the mean of the cth channel in the current batch; N represents the batch size;
[0123] 2. Normalization, specifically:
[0124] ,
[0125] where, represents the output result after normalization; represents the output result after global average pooling; μ c represents the mean of the c-th channel in the current batch; represents the variance of the c-th channel in the current batch;
[0126] 3. Affine transformation:
[0127] ,
[0128] where, represents the output result after normalization and affine transformation; γ c represents the scaling factor, used to adjust the distribution scale after normalization; β c represents the offset, used to recover the possible distribution offset after normalization; represents the output result after normalization.
[0129] Subsequently, through the fully connected layer (FC), whose output dimension is set to 1024, the features are further mapped and transformed to enrich the expression ability of the features.
[0130] Map the input vector to the target dimension D through the weight matrix:
[0131] ,
[0132] where, X represents the input matrix, which is the output after batch normalization processing, W represents the weight matrix, mapping the input from C dimensions to D dimensions; b represents the bias vector, and Z represents the output matrix, realizing the feature space transformation.
[0133] After that, perform the batch normalization (BN) operation again to further regularize the data distribution and make full preparations for the subsequent classification task; use the output of the fully connected layer as the input, and perform normalization for dimension D; calculate the mean μd and variance of each feature dimension d, and apply the affine transformation after normalization, specifically:
[0134] ,
[0135] where, represents the output; represents the scaling factor of dimension D; represents the original input value of the n-th sample and the d-th dimension, coming from the input of the fully connected layer; represents the mean of the d-th feature in the current batch; represents the variance of the d-th feature in the current batch; represents the offset of dimension D.
[0136] Finally, connect the classification layer to achieve the classification prediction of the target, specifically as follows:
[0137] The fully connected layer is mapped to the number of classes:
[0138] ,
[0139] where W cls represents the classification weight matrix; b cls represents the classification bias vector; Z represents the input matrix, with the output after the second batch normalization as the input; logitsk represents the unnormalized class score.
[0140] Softmax activation:
[0141] ,
[0142] where K represents the total number of classes; logits k represents the unnormalized class score (i.e., the output of the fully connected layer); is the probability value after Softmax, representing the probability that the sample belongs to the k-th class.
[0143] This unique structural design, namely the combination of GAP-BN-FC-BN, enables the improved ResNet50 model to exhibit distinct characteristics and advantages in feature extraction and classification performance compared to the traditional ResNet50 model.
[0144] To improve the model performance, an exploratory study was conducted on the CBAM (Convolutional Block Attention Module) attention mechanism. CBAM is only used on the output feature map of the last stage of ResNet50. During the model training process, in view of the uneven distribution of data of various classes in the training set, a series of targeted algorithms were actively adopted. For example, the class weighting algorithm assigns corresponding weights to different classes according to their proportions in the dataset, so as to pay more attention to the minority class samples during the loss calculation, enabling the model to more accurately learn the feature differences of various classes; the data sampling algorithm adjusts the relative quantities of data of various classes by oversampling the minority class samples or undersampling the majority class samples to achieve relative balance of data distribution; the EqualizationLoss v2 (Class Balanced Loss Function v2) algorithm starts from the loss function level and balances the contributions of different class samples to model training through special calculation methods; the BBN (Batch Balanced Network) algorithm starts from the perspective of batch data, dynamically adjusts and balances the data of different classes within each training batch to ensure that the model can fully learn the feature information of various class samples in each iteration, thereby enhancing the generalization ability and recognition accuracy of the model in the scenario of unbalanced data.
[0145] Longhorn beetles belonging to the same genus show a high degree of similarity in terms of color, shape, and markings, which poses a very serious challenge to the accurate identification of longhorn beetle species.
[0146] In this embodiment, the convolutional block attention module CBAM is introduced, which can guide the model to focus on the key information of the longhorn beetle itself and significantly enhance the feature extraction efficiency of the network. The CBAM attention mechanism includes two major parts: channel attention and spatial attention.
[0147] The calculation process of channel attention is as follows:
[0148] ,
[0149] where M c (F) represents the output of the channel attention module, that is, the weight map after applying channel attention to the input feature map F. represents the sigmoid activation function, which is used to compress the output value between 0 and 1 to generate the attention weight; MLP represents the multi-layer perceptron, AvgPool(F) represents global average pooling of the input feature map F; MaxPool(F) represents global maximum pooling of the input feature map F; W0 and W1 represent the weight matrices of two convolutional layers in the MLP; the first convolutional layer W0 is usually used for dimensionality reduction, and the second convolutional layer W1 is used to restore the number of channels or generate the final attention weight. represents the result after global average pooling of each channel of the feature map F; represents the result after global maximum pooling of each channel of the feature map F.
[0150] The calculation formula for spatial attention is:
[0151] ,
[0152] where M s (F) represents the output of the spatial attention module, that is, the weight map after applying spatial attention to the input feature map F. σ represents the Sigmoid activation function, which is used to compress the output value between 0 and 1 to generate the attention weight. f 7×7 represents a 7×7 convolutional operation, which is used to perform convolution processing on the concatenated feature map. AvgPool(F) represents global average pooling of the input feature map F, MaxPool(F) represents global maximum pooling of the input feature map F; [AvgPool(F);MaxPool(F)] represents concatenating the results of global average pooling and global maximum pooling in the channel dimension; and respectively represent the feature maps after global average pooling and global max pooling, which are concatenated together as f 7×7 The input of the convolution operation.
[0153] In this embodiment, the CBAM attention mechanism is ingeniously incorporated after the deformable convolution. The feature map will sequentially undergo in-depth processing of channel attention and spatial attention, precisely learning the importance degree of each feature channel and spatial dimension. In this unique way, the model is prompted to focus on the specific regional information where the longhorn beetles are located, significantly enhancing the weight ratio of color and texture information, while effectively reducing the adverse interference of background noise, thereby strongly promoting the steady improvement of the model's recognition ability. CBAM belongs to a lightweight module. After embedding them into the improved ResNet50 model, it can not only significantly improve the accuracy of longhorn beetle recognition, but also will not cause a substantial increase in network parameters. While ensuring the improvement of model performance, it takes into account the lightweight and efficiency of the model.
[0154] The test environment and evaluation indicators in this embodiment are specifically as follows:
[0155] The test environment is a 64-bit Windows 11 operating system, the processor is an i7-12700H, an NVIDIA GeForce RTX4090 Laptop GPU, 32GB of running memory, and 24GB of video memory.
[0156] To accurately evaluate the performance of the longhorn beetle detection model, this embodiment uses indicators such as precision (P), recall (R), and average precision (Ap).
[0157] Precision mainly demonstrates the accuracy of the model in classifying samples. Recall focuses on reflecting the model's ability to discover positive samples, while average precision can comprehensively reflect the overall performance of the model in object detection. For the longhorn beetle target recognizer, the indicators used include Top1 accuracy (i.e., the first hit rate) and Top2, Top3 accuracies, etc. The number of model parameters is a key indicator to measure the size of the model. Given that the number of model parameters is often quite large, it is usually measured in units of M (million). The number of parameters of the network is mainly composed of the number of parameters of the convolutional layer and the fully connected layer. Generally speaking, the lower the number of parameters, the more lightweight the model is, and the more convenient it is to be deployed and applied in resource-constrained environments such as mobile devices. Through the comprehensive consideration of these indicators, the actual effectiveness and performance level of the longhorn beetle detection model can be understood more deeply and comprehensively, thereby providing solid and reliable data support and theoretical basis for subsequent model optimization and improvement and related research work.
[0158] In this embodiment, the result analysis includes:
[0159] 1. Evaluation of longhorn beetle detection and recognition performance
[0160] For the longhorn beetle detector model, it is tested by applying the model to a set of image datasets of known longhorn beetle species. Currently, it can support approximately 577 taxonomic units, with 62,611 images in the training set and 6,995 images in the test set; calculate the accuracy and recall rate of the model to evaluate its recognition effect. The experimental results show that the longhorn beetle detector model based on YOLOv5 performs excellently in terms of accuracy and recall rate. The precision P, recall rate R, and mean average precision mAP are 0.97824, 0.97724, and 0.98935 respectively.
[0161] The longhorn beetle object recognizer can currently support approximately 689 taxonomic units and more than 70,000 images. The dataset is divided into a training set and a test set in a ratio of 9:1, with 70,330 images in the training set and 7,854 images in the test set; on the internally constructed test set, the Top1 to Top5 and Top10 accuracies are 0.951, 0.982, 0.988, 0.991, 0.992, and 0.995 respectively.
[0162] 2. Comparison of the performance of recognition models under different attention mechanisms
[0163] In this embodiment, the CBAM attention mechanism is selected and applied to the longhorn beetle object recognizer for a comparative experiment, and the comparison results are shown in Table 1. The results show that different fusion strategies have a certain impact on the performance of the model. Baseline (that is, the original ResNet50 Backbone does GAP and then directly goes to the classification layer).
[0164] Table 1 Performance indicators of 4 object classification models
[0165]
[0166] To sum up, the longhorn beetle detection model constructed based on the YOLOv5 model in this application performs excellently in terms of precision, recall rate, and mean average precision. After comparing the performance of different models in the longhorn beetle recognition task through experiments, it is found that although the longhorn beetle recognition model integrated with the attention mechanism slightly increases in resource occupancy, it has a significant advantage in terms of accuracy. In actual application scenarios, the appropriate fusion strategy can be flexibly selected according to specific requirements, and appropriate parameter adjustment and optimization means can be adopted to build a more accurate and efficient longhorn beetle recognition model system, so as to provide more powerful technical support and guarantee for longhorn beetle recognition related work.
[0167] Lightweight deployment: Through the ONNX Runtime inference engine and the FastAPI framework, the model is integrated into the small program to support real-time recognition on the mobile side.
[0168] The model carried by the mini-program uses the ONNX Runtime inference engine and can run efficiently directly on the CPU. After the user opens the mini-program, they can upload the image to be recognized by taking a photo or selecting a picture from the album. Its inference process only takes about 400 milliseconds, and the overall request response time can also quickly complete the image processing and recognition process within about 1 second, and accurately present the information of the five species with the highest probabilities of the insect, including its Chinese name, scientific name, and confidence level.
[0169] From the actual results, when using this mini-program to carry out the longhorn beetle recognition task in detection scenarios such as longhorn beetle ecological pictures with extremely complex environmental backgrounds, the accuracy rate can be as high as over 91%; however, in the case of relatively simple image backgrounds, the accuracy rate is only 75%. Delving into the root cause of this phenomenon, it is initially speculated that it is due to the obvious lack of diversity in the training set. Specifically, most of the pictures in the training set are mainly natural scenes, lacking white-background pictures, thus resulting in differences in the recognition ability of longhorn beetle specimen images.
[0170] A longhorn beetle image detection and recognition method, according to the above longhorn beetle image detection and recognition system, includes the following steps:
[0171] S1: Obtain the required image data and annotate it;
[0172] S2: Use the longhorn beetle detector to detect the longhorn beetle target in the input image and perform target cropping;
[0173] S21: Data preprocessing;
[0174] Extract images from the collected image data as the sample set, and accurately annotate this sample set so that the drawn bounding box closely surrounds the boundary of the longhorn beetle target to locate the position of the longhorn beetle; then divide the entire data set into a training set and a validation set according to a ratio of 9:1.
[0175] S22: Train the longhorn beetle detector and detect the target area;
[0176] Use the longhorn beetle detector to detect the data set, expand the detection box into a square according to the detected position and then crop the image, so as to obtain the longhorn beetle target image and save it to the corresponding species folder.
[0177] Detect the longhorn beetle target area in the image through the YOLOv5s model.
[0178] S3: Use the longhorn beetle recognizer to recognize and classify the cropped and cleaned images;
[0179] S31: Automatically clean the longhorn beetle target image data set to obtain a relatively clean data set;
[0180] S32: Input the target area into the improved ResNet50 model, and combine with the CBAM attention mechanism to output the species classification result;
[0181] S4: Implement image uploading, real-time inference, and result display in the applet.
[0182] The applet in this embodiment focuses on the specific field of longhorn beetles. By deeply analyzing the unique morphological characteristics, texture details, and ecological habits of longhorn beetles, and through continuously optimized algorithm models, it can more accurately identify different species of longhorn beetles. The prior art developed an insect recognition model based on YOLOv5 + CABM, mainly for the recognition of general insects. The model performs well in indicators such as accuracy, recall rate, and mAP@0.5. Compared with YOLOv5l, the accuracy of the improved model has increased by 1%, the recall rate has increased by 1.3%, the mAP@0.5 value has increased by 1.7%, and the F1 score has increased by 0.02. In contrast, this embodiment uses the YOLOv5 detection model and the ResNe50 + CABM recognition model, focusing on longhorn beetle recognition. The dataset is larger in scale, the performance indicators are better, and the specific values are clear, reflecting good performance in the longhorn beetle recognition task. At the same time, a mobile applet for longhorn beetle recognition at ports has been developed, targeting the port inspection and quarantine scenario, which facilitates the staff to quickly and accurately identify the species of longhorn beetles at the port site, improves the inspection and quarantine efficiency, and the application scenario is more targeted and closely integrated with the mobile terminal.
[0183] The present invention conducts training and testing on more than 60,000 longhorn beetle image samples of 576 species of longhorn beetles, deeply analyzes and compares the performance of the two models in the tasks of longhorn beetle image detection and recognition, aiming to provide strong technical support and data basis for ecological research related to longhorn beetles, forestry pest control and other fields, and promotes the further development and application of longhorn beetle image intelligent analysis technology.
[0184] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A longhorn beetle image detection and recognition system, characterized in that, Including: Data preprocessing module: Collect the required data and label the collected data set. Longhorn beetle detection module: Based on the YOLOv5s model, accurately locate the position of the longhorn beetle in the image. Data cropping module: Enlarge the detection box according to the detected position of the longhorn beetle and then crop the image to obtain a data set of longhorn beetle target images. Longhorn beetle recognition module: Based on the improved ResNet50 model, incorporating the CBAM attention mechanism, for high-precision recognition of longhorn beetle species. Lightweight deployment: Integrate the model into the applet through the ONNX Runtime inference engine and the FastAPI framework to support real-time recognition on mobile devices.
2. The longhorn beetle image detection and recognition system according to claim 1, characterized in that, The specific content of the data preprocessing module is: Data collection: Collect specimen images and ecological images including longhorn beetles. Data preprocessing: Randomly extract a certain number of images from the collected image library as a sample set, use the LabelMe tool to label the position of the longhorn beetle in the image, and convert the labeled data set into the YOLOv5 format; finally, divide the data set into a training set and a validation set according to a ratio.
3. The longhorn beetle image detection and recognition system according to claim 1, characterized in that The specific content of the longhorn beetle detection module is: Based on the YOLOv5 model architecture, use two longhorn beetle detectors. One detector selects the YOLOv5l model structure to label the position of the longhorn beetle in the data set; the other detector needs to select the YOLOv5s model structure to achieve fast longhorn beetle detection. During the training process, a comprehensive evaluation is performed after each round of training, and the model generated in the current round is saved synchronously.
4. The longhorn beetle image detection and recognition system according to claim 1, characterized in that, The specific content of the data cropping module is: Data cropping: Use the previous longhorn beetle detector to detect each image in the entire data set, enlarge the detection box to a square according to the detected position and then crop the image, so as to obtain longhorn beetle target images and save them to the corresponding species folders. Finally, a data set of longhorn beetle target images is formed. Data cleaning: Automatically clean the data set of longhorn beetle target images using the Cleanlab library.
5. The longhorn beetle image detection and recognition system according to claim 1, characterized in that The specific content of the longhorn beetle recognition module is: The improved ResNet50 model includes network structure adjustment, attention mechanism fusion, and classification layer optimization. Specifically: Network structure adjustment: Change the stride of the last stage in the ResNet50 module to 1 to retain more refined feature maps. Attention mechanism fusion: Embed the CBAM module after the deformable convolution, and focus on the key texture and color features of the longhorn beetle through channel and spatial attention weighting. Classification layer optimization: Use the GAP-BN-FC-BN structure to replace the traditional fully connected layer, combined with the label-smoothing cross-entropy loss function to alleviate class imbalance.
6. The longhorn beetle image detection and recognition system according to claim 5, wherein, The CBAM attention mechanism includes channel attention and spatial attention; the feature map undergoes depth processing of channel attention and spatial attention in sequence to learn the importance degree of each feature channel and spatial dimension. The calculation process of channel attention is: , Among them, M c (F) represents the output of the channel attention module, represents the Sigmoid activation function; MLP represents a multi-layer perceptron, AvgPool(F) represents global average pooling of the input feature map F; MaxPool(F) represents global max pooling of the input feature map F; W0 and W1 represent the weight matrices of two convolutional layers in the MLP; represents the result after global average pooling for each channel of the feature map F; represents the result after global max pooling for each channel of the feature map F.
7. The longhorn beetle image detection and recognition system according to claim 6, characterized in that, The calculation formula of spatial attention is: , Among them, M s (F) represents the output of the spatial attention module, σ represents the Sigmoid activation function, and f 7×7 represents a 7×7 convolution operation, AvgPool(F) represents global average pooling on the input feature map F, and MaxPool(F) represents global max pooling on the input feature map F; [AvgPool(F); MaxPool(F)] represents concatenating the results of global average pooling and global max pooling along the channel dimension; and respectively represent the feature maps after global average pooling and global max pooling.
8. The longhorn beetle image detection and recognition system according to claim 5, characterized in that, The output result of the CBAM attention mechanism will first go through a global average pooling operation: By taking the average of the entire feature map in the spatial dimension, the feature map is converted into a feature vector with global information. The specific calculation formula is: , Among them, represents the output of the global average pooling operation on the feature map output by CBAM, H represents the height of the feature map, W represents the width of the feature map; X represents the input feature map; i, j represent the spatial coordinates of the feature map; n represents the nth sample in the batch; c represents the cth channel of the feature map.
9. The longhorn beetle image detection and recognition system according to claim 8, wherein, Perform batch normalization after global average pooling. The specific process is as follows: Calculate batch statistics, including calculating the mean and variance; Calculate the mean, channel by channel. The specific calculation formula is: , Among them, μ c represents the mean value of the c-th channel in the current batch; N represents the batch size; represents the output data of global average pooling; Calculate the variance. The specific calculation formula is: , Among them, represents the variance of the c-th channel in the current batch, represents the smoothing term; represents the output data of global average pooling; μ c represents the mean of the c-th channel in the current batch; N represents the batch size; Normalize, specifically: , Among them, represents the output result after normalization; represents the output result after global average pooling; μ c represents the mean value of the c-th channel in the current batch; represents the variance of the c-th channel in the current batch; Affine transformation, specifically: , Among them, represents the output result after normalization and affine transformation; γ c represents the scaling factor; β c represents the offset; represents the output result after normalization.
10. A longicorn beetle image detection and recognition method, according to the longicorn beetle image detection and recognition system described in any one of claims 1-9, characterized in that: Include the following steps: S1: Obtain the required image data and annotate it; S2: Use the longhorn beetle detector to perform longhorn beetle target detection on the input image and perform target cropping; Extract images from the collected image data as a sample set and accurately annotate the sample set so that the drawn bounding box tightly surrounds the boundary of the longhorn beetle target to locate the position of the longhorn beetle; Use the longhorn beetle detector to detect the data set, expand the detection box into a square according to the detected position and then crop the image, so as to obtain the longhorn beetle target image and save it to the corresponding species folder; Detect the longhorn beetle target area in the image through the YOLOv5s model; S3: Use the longhorn beetle recognizer to perform recognition and classification on the cropped and cleaned images; S31: Automatically clean the longhorn beetle target image data set to obtain a relatively clean data set; S32: Input the target area into the improved ResNet50 model, and combine the CBAM attention mechanism to output the species classification result; S4: Implement image upload, real-time inference and result display in the applet.
Citation Information
Cited By
Contraband detection method based on double-view-angle X-ray image fusion
CN121213888A