Allium chinense key point identification method and device, program product and storage medium

By improving the C3k2 module of the YOLO11 network to the C3k2_RepVGG module and the C2PSA module to the dual-layer routing attention mechanism module, and inserting the feature enhancement layer MobileNet Variants into the neck-connected network, the problems of insufficient long-distance feature dependence and low feature fusion efficiency of the YOLOv11-Pose baseline model in the detection of key points of the head are solved, and higher detection accuracy and robustness are achieved.

CN120472208AActive Publication Date: 2025-08-12HUAZHONG AGRI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510510559.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-12
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing YOLOv11-Pose baseline model has problems such as insufficient long-distance feature dependence, low feature fusion efficiency, poor robustness and insufficient generalization ability when detecting key points of the head, resulting in a decrease in detection accuracy and large grading errors in complex environments.

Method used

The improved YOLO11 network is adopted, and the C3k2 module in the backbone network is replaced as the C3k2_RepVGG module, the C2PSA module is replaced as the dual-layer routing attention mechanism module, and the feature enhancement layer MobileNet Variants are inserted into the neck-connected network to enhance feature extraction and information exchange capabilities.

Benefits of technology

It improves the critical point detection accuracy and robustness of the model in complex environments, reduces the computational complexity, enhances the ability to capture long-distance dependencies, and improves the generalization ability and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472208A_ABST
    Figure CN120472208A_ABST
Patent Text Reader

Abstract

The invention discloses a method for identifying key points of allium chinensis, which is characterized in that an improved YOLOv11 network adopted by the method is improved on the basis of a YOLOv11-Pose baseline model, a C2PSA module in a backbone network of the YOLOv11-Pose baseline model is replaced by a double-layer routing attention mechanism module, and long-distance feature perception is enhanced through coarse-grained routing screening and fine-grained Token-to-Token attention; a feature enhancement layer MobileNet Variants is inserted between each splicing module of a neck connection network of the YOLOv11-Pose baseline model and a C3k2 module of an adjacent next layer, and a depth separable convolution and a Leaky ReLU activation function are adopted to enhance the feature learning ability; a C3k2 module in a backbone network of a YOLOv11-Pose baseline model is replaced by a C3k2RepVGG module, and channel shuffling and re-parameterization technologies are fused, so that the information exchange capability between channels is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural image processing, and in particular relates to a method for identifying key points of radish, and also relates to computer equipment, program products and storage media, which are suitable for automatic identification of key points of radish. Background Art

[0002] With the widespread application of deep learning technology in agricultural automation, vision-based object detection and keypoint recognition algorithms have become a research hotspot. As a root crop with an irregular shape and large size variation, radish faces numerous challenges in intelligent harvesting and grading using traditional mechanical grippers. For example, inaccurate gripping points can easily lead to damage. Existing methods that rely on manual experience or traditional image processing techniques suffer from poor robustness and generalization capabilities. While the YOLO series of models have demonstrated certain advantages in agricultural object detection, research on keypoint detection and grading for radish, a type of crop with an irregular shape, is still in its infancy.

[0003] The YOLO series of models is widely used in agricultural inspection. YOLOv8-Pose, among others, can detect key points and has been applied in scenarios such as strawberry stalk positioning. However, it is not optimized for long-range feature dependencies and small objects, making it prone to missed detections in complex backgrounds. YOLOv11-Pose, the latest iteration of the model, introduces the C3k2 module and the C2PSA attention mechanism. While this offers some improvements, C2PSA lacks a dynamic filtering mechanism, limiting its feature extraction capabilities in complex environments. Other methods, such as traditional machine vision's Fourier shape classification, struggle to adapt to the irregular shapes of radishes. While the lightweight improved model reduces the number of parameters, it is ineffective in detecting densely packed radishes.

[0004] The YOLOv11-Pose baseline model consists of a backbone, neck, and head. Its workflow involves extracting multi-scale feature maps from the input image through the backbone. The neck fuses these features and passes them to the head, which then uses the C2PSA module to output keypoints and detection bounding boxes. However, it suffers from issues such as insufficient capture of long-range dependencies, low feature fusion efficiency, and insufficient lightweightness. Core issues with existing technologies include weak dynamic perception, insufficient interaction between multi-scale features, and a reliance on single-feature classification methods. These issues lead to reduced detection accuracy in complex environments and large classification errors.

[0005] (1) The existing YOLOv11-pose baseline model uses the C2PSA attention mechanism, which has the disadvantage of low performance in detecting the key points of garlic sprouts;

[0006] (2) The existing YOLOv11-pose baseline model uses the C3k2 module, which has the disadvantages of low information exchange ability and feature extraction ability between different channel features;

[0007] (3) The existing YOLOv11-pose baseline model has the disadvantages of poor model robustness and low generalization ability. Summary of the Invention

[0008] The purpose of the present invention is to address the above-mentioned problems existing in the prior art and provide a method for identifying key points of radish heads, as well as a computer device, a program product and a storage medium. The improved YOLO11 network adopted in the present invention is improved based on the original YOLOv11-Pose baseline model. It can not only identify the key points of radish heads in normal postures, but also has a good detection effect on the key points of radish heads in complex situations such as camera overexposure and mutual covering of radishes, thereby providing a method for identifying key points of radish heads with high recognition accuracy.

[0009] The above-mentioned purpose of the present invention is achieved by the following technical means:

[0010] A method for identifying key points of scallion heads comprises the following steps:

[0011] Step 1: Obtain multiple scallion images, and obtain scallion key points for each scallion image. The scallion images and corresponding scallion key points are input samples and labels, respectively. Each input sample and corresponding label is a sample. All samples constitute a sample set, and the sample set is divided into a training set and a test set according to the proportion.

[0012] Step 2: Build and improve the YOLO11 network;

[0013] Step 3: Construct the loss function Loss of the improved YOLO11 network;

[0014] Step 4: Input the training set into the improved YOLO11 network for training, and save the model parameters after training is completed;

[0015] Step 5: Input the test set into the improved YOLO11 network and output the key point prediction results corresponding to each input sample in the test set;

[0016] Step 6: Obtain the scallion head image to be identified, input the scallion head image to be identified into the improved YOLO11 network, and output the key point prediction result corresponding to the scallion head image to be identified.

[0017] The improved YOLO11 network mentioned above includes the following improvements based on the YOLOv11-Pose baseline model:

[0018] Improvement 1: Improve the YOLO11 network by replacing the C3k2 module in the backbone network of the YOLOv11-Pose baseline model with the C3k2_RepVGG module;

[0019] Improvement 2: Improve the YOLO11 network by replacing the C2PSA module in the backbone network of the YOLOv11-Pose baseline model with a two-layer routing attention mechanism module;

[0020] Improvement 3: Improve the YOLO11 network by inserting feature enhancement layers, MobileNet Variants, between each splicing module of the neck connection network of the YOLOv11-Pose baseline model and the C3k2 module of the adjacent next layer.

[0021] As mentioned above, the C3k2_RepVGG module includes a channel splitting module, a reparameterized convolution module, a splicing module, a bottleneck layer, and a channel shuffling module. The feature map input to the C3k2_RepVGG module first passes through the channel splitting module. The channel splitting module divides the input feature map into two sub-feature maps with half the number of channels. One of the sub-feature maps is input to the splicing module of the C3k2_RepVGG module.

[0022] When C3k is True, the other sub-feature map is input into the bottleneck layer, which performs deep feature extraction on the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; when C3k is False, the other sub-feature map is input into the RepVGG module, which reparameterizes the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; finally, the splicing module of the C3k2_RepVGG module splices and fuses the two input sub-feature maps and outputs them to the channel shuffling module for grouping, shuffling, and reorganization, and finally outputs the reorganized feature map.

[0023] As mentioned above, the RepVGG module adopts a multi-branch structure in the training phase and a single-branch structure in the inference phase;

[0024] The multi-branch structure of the RepVGG module includes a 1x1 convolution layer, a 3x3 convolution layer, an identity mapping, and a SilU activation function. The sub-feature maps input to the RepVGG module are respectively subjected to a 1x1 convolution layer, a 3x3 convolution layer, and an identity mapping before being added and fused. The fused feature maps are then subjected to a SilU activation function, which then outputs the result to the concatenation module of the C3k2_RepVGG module.

[0025] The single-branch structure adopted by the RepVGG module in the inference stage includes a 3x3 convolution layer and a SilU activation function. The sub-feature map input to the RepVGG module passes through the 3x3 convolution layer and the SilU activation function in turn and is output to the splicing module of the C3k2_RepVGG module.

[0026] As mentioned above, the dual-layer routing attention mechanism module takes the input feature map The following processing steps are included:

[0027] First, the input feature map X is divided into non-overlapping regions, each of which consists of feature vectors, reshape the input feature X into the input feature , And obtain the query tensor corresponding to query, key, and value through linear projection , key tensor , and value tensors , query tensor , key tensor , and value tensors Calculated based on the following formula:

[0028]

[0029] in, is a set of real numbers, H is the height of the feature map X, W is the width of the feature map X, C is the number of channels of the feature map X, are the projection weights of query, key, and value respectively;

[0030] Secondly, the query tensor Q is averaged region by region to obtain the matrix , key tensor After averaging each region, the matrix is obtained , we get the matrix by matrix multiplication and matrix The inter-region adjacency matrix ,Using directed graph method to build inter-region routing,inter-region adjacency matrix Calculated based on the following formula:

[0031]

[0032] Where, for The transposed matrix of

[0033] Again, take the first k most relevant associations in each row to prune the directed graph and build a routing index matrix , routing index matrix Calculated based on the following formula:

[0034]

[0035] Where, is the adjacency matrix between regions Keep the associations of the top k maximum values in each row row by row;

[0036] In the routing index matrix On the top, the token-to-token attention mechanism is applied. The token-to-token attention mechanism starts with the routing index matrix Collect key tensors Sum tensor , the key tensor Sum tensor Calculated based on the following formulas:

[0037]

[0038] Where, gather is the collection operation from the routing index matrix;

[0039] Output of the token-to-token attention mechanism Calculated based on the following formula:

[0040]

[0041]

[0042]

[0043] Among them, LCE(V) is the local context enhancement term, which uses the depth convolution pair tensor with a convolution kernel size of 5×5 Parameterize; for The dimension size, for The transposed matrix of That is , for The i-th element in, i ranges from 1 to L, L is length.

[0044] As described above, the feature enhancement layer MobileNet Variants includes a depth convolution module, a first activation function layer, a first residual connection layer, a point-by-point convolution module, a second activation function layer, and a second residual connection layer connected in sequence;

[0045] The first and second activation function layers are based on the following formula:

[0046]

[0047] Where, is the first activation function layer or the second activation function, is the feature map input to the first activation function layer or the second activation function, and α is a constant.

[0048] As mentioned above, the loss function Loss of the improved YOLO11 network is calculated based on the following formula:

[0049]

[0050] Where, is the intersection-over-union loss function, is the mean square error loss function, is the weight coefficient of the intersection-over-union loss function, is the weight coefficient of the mean square error loss function.

[0051] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for identifying key points of radish heads when executing the computer program.

[0052] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for identifying key points of radish heads.

[0053] A computer program product includes a computer program, which implements the steps of the onion key point identification method described above when the computer program is executed by a processor.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] (1) The dual-layer routing attention mechanism module BRA adopted in the present invention enhances the model's ability to capture long-distance dependencies by dynamically identifying regions with higher relevance and giving them higher attention. In addition, the calculation only involves GPU-friendly dense matrix multiplication, thereby improving the model's feature extraction capability and computational efficiency.

[0056] (2) The present invention proposes a C3k2_RepVGG module that integrates the feature channel shuffling and reparameterized convolution module on the C3k2 module proposed in the original YOLO11. The reparameterization of the RepVGG structure speeds up the model inference time, and the shuffling operation between the feature channels of ShuffleNet improves the fluidity between the model feature information, realizing information exchange between different channel features while maintaining a low computational complexity of the model.

[0057] (3) The present invention adopts the MobileNet Variants feature enhancement layer, which extracts features from different channels by using deep convolution and activation function LeakyReLU on the input features, and then uses point-by-point convolution and activation function LeakyReLU to realize information exchange between channels. The residual connection is used to fuse the feature map of the original input with the results of the two convolutions respectively, which effectively prevents the problem of neuron death during training, making the model more robust and generalizable. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Schematic diagram of the structure of the improved YOLO11 network of the present invention (wherein SPPF is Spatial Pyramid Pooling-Fast (Spatial Pyramid Pooling-Fast), and Upsample is the upsampling module);

[0059] Figure 2 Schematic diagram of the structure of the two-layer routing attention mechanism (where mm is the matrix multiplication operation and gather is the collection operation from the routing index matrix);

[0060] Figure 3 Schematic diagram of the structure of the C3k2_RepVGG module of the present invention;

[0061] Figure 4 This is a schematic diagram of the structure of the feature enhancement layer MobileNet Variants of the present invention;

[0062] Figure 5 Schematic diagram of the key points of scallion of the present invention. DETAILED DESCRIPTION

[0063] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below with reference to the embodiments. The embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0064] Example 1:

[0065] A method for identifying key points of scallion heads comprises the following steps:

[0066] Step 1: Obtain multiple scallion images to form a sample set. Each sample in the sample set includes an input sample and a label. The sample set is divided into a training set and a test set in proportion. Specifically, the following steps are included:

[0067] Step 1.1: Use a distortion-free camera to photograph scallions to obtain multiple scallion images. For each scallion image, obtain scallion key points. In this embodiment, the scallion key points are the vertices and two transverse coordinates of the scallion. Each scallion image is used as an input sample. The scallion key points of each input sample serve as the label of the corresponding input sample. Each input sample and its corresponding label are considered as a sample, and all samples constitute a sample set.

[0068] Step 1.2: Expand the sample set by data augmentation (such as mosaic stitching, HSV color perturbation, and random rotation (±30°)) to obtain an expanded sample set to improve the generalization ability of the model of the present invention;

[0069] Step 1.2: Divide the expanded sample set into training set and test set in proportion;

[0070] Step 1.3: Divide the sample set into training set and test set.

[0071] Step 2: Build an improved YOLO11 network. The improved YOLO11 network includes a backbone network (Backbone), a neck connection network (Neck), and a detection head (Head). Compared with the original YOLOv11-Pose baseline model, the improved YOLO11 network of the present invention includes the following improvements:

[0072] Improvement 1: The improved YOLO11 network of the present invention replaces the C3k2 module in the backbone network of the original YOLOv11-Pose baseline model with the C3k2_RepVGG module (the C3k2_RepVGG module is a module that integrates channel shuffling and reparameterized convolution, which can improve feature fusion efficiency);

[0073] The C3k2_RepVGG module includes a channel splitting module (Channel Split), a reparameterized convolution module (RepVGG), a splicing module (Concat), a bottleneck layer (BottleNeck), and a channel shuffling module (ChannelShuffle). The feature map input to the C3k2_RepVGG module first passes through the channel splitting module. The channel splitting module divides the input feature map into two sub-feature maps with half the number of channels. One of the sub-feature maps is input to the splicing module of the C3k2_RepVGG module; when C3k is True, the other sub-feature map is input to the bottleneck layer. The bottleneck The layer performs deep feature extraction on the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; when C3k is False, the other sub-feature map is input into the RepVGG module, and the RepVGG module reconstructs the structural parameters (i.e., reparameterizes) of the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; finally, the splicing module of the C3k2_RepVGG module splices and fuses the two input sub-feature maps and outputs them to the channel shuffling module for grouping, shuffling and reorganization, and finally outputs the reorganized feature map.

[0074] The RepVGG module uses a multi-branch structure during the training phase (Train), which includes convolution kernel branches of different sizes (such as 3x3 convolution, 1x1 convolution, and identity mapping) and the SilU activation function. During the inference phase (Predict, which uses the model trained on the training set to infer the test set), it is simplified to a single-branch structure through reparameterization, including only 3x3 convolution and the SilU activation function.

[0075] As an implementable embodiment, the multi-branch structure of the RepVGG module includes a 1x1 convolution layer, a 3x3 convolution layer, an identity mapping, and a SilU activation function. The sub-feature maps input to the RepVGG module are respectively subjected to the 1x1 convolution layer, the 3x3 convolution layer, and the identity mapping and then added and fused. The feature maps after addition and fusion are passed through the SilU activation function, and the SilU activation function inputs the output result into the splicing module of the C3k2_RepVGG module.

[0076] As an implementable method, the single-branch structure adopted by the RepVGG module in the training stage includes a 3x3 convolution layer and a SilU activation function. The sub-feature map input to the RepVGG module passes through the 3x3 convolution layer and the SilU activation function in turn and is output to the splicing module of the C3k2_RepVGG module.

[0077] The C3k2_RepVGG module of the present invention integrates the RepVGG structure and the ShuffleNet structure. The RepVGG structure adopts a multi-branch structure in the training phase by utilizing reparameterization, and is simplified to a single-branch structure in the inference phase by reparameterization. The ShuffleNet structure shuffles the feature channels. When C3k is True, the C3k2_RepVGG module uses multi-layer BottleNeck to extract deep features from the input features and splices them with the directly transmitted features. When C3k is False, the BottleNeck module is replaced by the RepVGG module.

[0078] Improvement 2: The improved YOLO11 network of the present invention replaces the C2PSA module in the backbone network of the original YOLOv11-Pose baseline model with a dual-layer routing attention mechanism module (BRA module, a dynamic sparse attention mechanism that enhances the model's perception of long-distance features). The dual-layer routing attention mechanism based on the visual Transformer filters out most irrelevant key-value pairs in coarse-grained areas, retaining a small number of key-value pairs with strong correlation, and uses a token-to-token attention mechanism in fine-grained areas to capture the key features of the target and enhance long-distance feature perception. Specifically:

[0079] For a feature map input to a two-layer routing attention mechanism module, , H is the height of feature map X, W is the width of feature map X, C is the number of channels of feature map X, first divide the input feature X into non-overlapping regions, each of which consists of feature vectors, reshape the input feature X into the input feature And obtain the query tensor corresponding to query, key, and value through linear projection , key tensor , and value tensors , query tensor , key tensor , and value tensors Calculated by the following formula:

[0080] (1)

[0081] in, is the set of real numbers, are the projection weights of query, key, and value respectively;

[0082] Then the query tensor Q is averaged region by region to obtain the matrix , key tensor After averaging each region, the matrix is obtained , we get the matrix by matrix multiplication and matrix The inter-region adjacency matrix of ,Using directed graph method to build inter-region routing,inter-region adjacency matrix Calculated by the following formula:

[0083] (2)

[0084] Where, for The transposed matrix of

[0085] The average of each region is to calculate the average value of Q and K of each region one by one for S×S non-overlapping regions. and .

[0086] Then, at the coarse-grained level, the directed graph is pruned by taking the top k most relevant associations in each row to construct a routing index matrix , routing index matrix The calculation formula is:

[0087] (3)

[0088] Where, Refers to finding the index of the first k maximum (or minimum) values from an array or matrix. This embodiment is for the inter-region adjacency matrix Keep the top k largest values in each row.

[0089] In the routing index matrix On the other hand, a fine-grained token-to-token attention mechanism is applied: the token-to-token attention mechanism starts with the routing index matrix Collect key tensors Sum tensor , from the routing index matrix The key tensor and value tensor are collected based on the following formulas:

[0090] (4)

[0091] Where, gather is the collection operation from the routing index matrix;

[0092] The output of the token-to-token attention mechanism is As shown in (5):

[0093] (5)

[0094] (6)

[0095] (7)

[0096] Among them, LCE(V) is the local context enhancement term, which uses the depth convolution pair tensor with a convolution kernel size of 5×5 Parameterize; for The dimension size, for The transposed matrix of That is , for The i-th element in, i ranges from 1 to L, L is length.

[0097] Improvement 3: The improved YOLO11 network of the present invention inserts a feature enhancement layer (MobileNet Variants, a lightweight convolutional neural network variant module) between each splicing module of the neck connection network of the original YOLOv11-Pose baseline model and the C3k2 module of the adjacent next layer. Based on the feature enhancement layer of depthwise separable convolution, it optimizes the model's lightweight and feature extraction capabilities. Specifically:

[0098] The feature enhancement layer MobileNet Variants includes a sequentially connected deep convolution module (DWConv, Depthwise Convolution), a first activation function layer (LeakyReLU), a first residual connection layer, a pointwise convolution module (PWConv, Pointwise Convolution), a second activation function layer, and a second residual connection layer. The feature map input to the feature enhancement layer MobileNet Variants passes through the deep convolution module and the first activation function layer in sequence. The deep convolution module and the first activation function layer extract features of different channels of the input feature map. The first activation function layer outputs the feature map after feature extraction to the first residual connection layer. The first residual connection layer fuses the input feature map of the lightweight convolution neural network variant module with the feature map output by the first activation function layer for the first time, and outputs the first fused feature map to the pointwise convolution module. Pointwise convolution and the second activation function layer are then used to realize inter-channel information exchange. The feature map output by the second activation function layer is input to the second residual connection layer. The second residual connection layer fuses the input feature map of the lightweight convolution neural network variant module with the feature map output by the second activation function layer for the second time, and outputs the second fused feature map.

[0099] Among them, the first activation function layer and the second activation function layer are based on the following formula, and the loss function Loss of the improved YOLO11 network is calculated based on the following formula:

[0100] (8)

[0101] Where, is the input feature map, α is a constant, and in this embodiment, the value is 0.01.

[0102] The lightweight convolutional neural network variant module replaces the ReLu activation function in the original MobileNet V1 model (a lightweight convolutional neural network) with the LeakyReLU activation function. The input feature map first extracts features from different channels through deep convolution and activation functions, then uses point-by-point convolution and activation functions to achieve information exchange between channels, and uses residual connections to fuse the original input feature map with the two convolution results.

[0103] The workflow of the network model of the present invention is:

[0104] The improved YOLO11n-Pose network is based on the YOLOv11-Pose baseline model. Its detailed data flow is as follows: Allium stalk images with a resolution of 640×640 pixels and containing various soil environments, lighting conditions, and occlusion scenarios, along with annotation data (including the vertices and the endpoints of the transverse diameter), are fed into the backbone network through two convolutional neural networks. The network then passes through the C3k2_RepVGG module (which integrates channel shuffling and reparameterized convolution, an improvement on the original YOLO11 C3k2 module) to receive the feature map from the previous layer. The training phase utilizes a multi-branch approach, firstly using a convolutional neural network layer combined with the C3k2_RepVGG reparameterization technique to optimize training efficiency, then feeding into the next convolutional neural network layer, which outputs multi-scale feature maps. The SPPF module receives deep features, extracts global context through a cascade of pooling layers, and outputs a fused feature map. Next is the neck connection network. A two-layer routing attention mechanism based on the visual Transformer receives the multi-scale features output by the backbone. It first performs coarse-grained routing to partition the feature map and calculate the inter-region adjacency matrix. It then performs fine-grained attention calculations, combining local context enhancement within the filtered region to output a feature map with enhanced long-range dependencies. The concat module receives the BRA output features and the shallow features output by the C3k2_RepVGG module. It fuses the multi-scale information through upsampling and concatenation, and outputs the neck fused features. The output of the C3k2_RepVGG module is then transferred to the concat module and enters the neck portion. The lightweight MobileNet Variants layer receives the neck fused features, extracts local features within the channel, fuses inter-channel information through depthwise convolution, replaces the activation function with LeakyReLU, and uses residual connections to retain the original input features. The enhanced feature map is then output. The prediction layer receives the enhanced feature map, predicts the key points of the enhanced feature map, and outputs it.

[0105] Step 3: Construct the loss function Loss of the improved YOLO11 network:

[0106] (9)

[0107] in, is the intersection-over-union loss function, is the mean square error loss function, is the weight coefficient of the intersection-over-union loss function, for The weight coefficient, in this embodiment, the weight ratio ( and The ratio is 3:1.

[0108] Step 4: Input the training set into the improved YOLO11 network for training, and save the model parameters after training is completed;

[0109] Step 5: Simplify the RepVGG structure of the trained improved YOLO11 network into a single structure to obtain the improved YOLO11 network in the inference stage. Input the test set into the improved YOLO11 network in the inference stage and output the key point prediction results corresponding to each input sample in the test set.

[0110] Step 6: Obtain the scallion head image to be identified, input the scallion head image to be identified into the improved YOLO11 network in the inference stage, and output the key point prediction result corresponding to the scallion head image to be identified.

[0111] Figure 5 The key point prediction results of a scallion head image are given, where the scallion head vertex U and the two horizontal diameter coordinates L and R are the prediction results of the improved YOLO11 network when the scallion head image to be identified is input into the inference stage. M and Z are the key points calculated based on the predicted scallion head vertex U and the two horizontal diameter coordinates L and R (that is, the scallion head detection frame is determined based on the identified scallion head vertex U and the two horizontal diameter coordinates L and R, point Z represents the intersection of the diagonals of the detection frame, and point M represents the midpoint of the line connecting points L and R).

[0112] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0113] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0114] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0115] Example 2:

[0116] This example collected images of large-leaf radish (a radish variety) from Chongyang County, Xianning City, Hubei Province, using an industrial camera with a resolution of 1920×1080. The images covered six different scenarios, including sunny, cloudy, obscured, and overexposed conditions. A total of 12,500 images were annotated, with key points including the radish tip, root, and bend. Data augmentation employed mosaic stitching, HSV color perturbation, and random rotation (±30°) to improve model generalization.

[0117] Training parameters: The model was trained on an NVIDIA RTX 3090 GPU with an initial learning rate of 0.01, a cosine annealing strategy, a batch size of 32, and a loss function that was a weighted combination of CIoU Loss (detection box) and MSE Loss (key points) (weight ratio 3:1).

[0118] Experimental results: The improved BMR-YOLO11-Pose achieves a box precision (Box-P) of 90.2% on the test set (a 2.4% improvement over the baseline), a keypoint mean error (MPE) of 2.8 pixels (a 18% reduction), a model size of only 5.91 MB, and an inference speed of 58.95 FPS.

[0119] A radish intelligent harvesting system was deployed on a farm in Chongyang County. The robotic arm adaptively adjusted its gripping angle based on keypoint coordinates, reducing the harvesting damage rate from 12% with traditional methods to 3.5% and improving grading accuracy to 81.72% (based on a CART decision tree). In an overexposure scenario test, under strong light conditions (light intensity >10,000 lux), the model used the BRA mechanism to suppress interference from bright areas, achieving a keypoint miss rate of only 5.1% (compared to 18.3% for the original YOLOv11-Pose baseline model).

[0120] It should be noted that the embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.

Claims

1. A method for identifying key points of scallion, characterized in that: The following steps are involved: Step 1: Obtain multiple scallion images, and obtain scallion key points for each scallion image. The scallion images and corresponding scallion key points are input samples and labels, respectively. Each input sample and corresponding label is a sample. All samples constitute a sample set, and the sample set is divided into a training set and a test set according to the proportion. Step 2: Build and improve the YOLO11 network; Step 3: Construct the loss function Loss of the improved YOLO11 network; Step 4: Input the training set into the improved YOLO11 network for training, and save the model parameters after training is completed; Step 5: Input the test set into the improved YOLO11 network and output the key point prediction results corresponding to each input sample in the test set; Step 6: Obtain the scallion head image to be identified, input the scallion head image to be identified into the improved YOLO11 network, and output the key point prediction result corresponding to the scallion head image to be identified.

2. The method for identifying key points of scallion according to claim 1, wherein: The improved YOLO11 network includes the following improvements based on the YOLOv11-Pose baseline model: Improvement 1: Improve the YOLO11 network by replacing the C3k2 module in the backbone network of the YOLOv11-Pose baseline model with the C3k2_RepVGG module; Improvement 2: Improve the YOLO11 network by replacing the C2PSA module in the backbone network of the YOLOv11-Pose baseline model with a two-layer routing attention mechanism module; Improvement 3: Improve the YOLO11 network by inserting feature enhancement layers, MobileNet Variants, between each splicing module of the neck connection network of the YOLOv11-Pose baseline model and the C3k2 module of the adjacent next layer.

3. The method for identifying key points of scallion according to claim 2, wherein: The C3k2_RepVGG module includes a channel splitting module, a reparameterized convolution module, a splicing module, a bottleneck layer, and a channel shuffling module. The feature map input to the C3k2_RepVGG module first passes through the channel splitting module. The channel splitting module evenly divides the input feature map into two sub-feature maps with half the number of channels, one of which is input into the splicing module of the C3k2_RepVGG module. When C3k is True, the other sub-feature map is input into the bottleneck layer, which performs deep feature extraction on the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; when C3k is False, the other sub-feature map is input into the RepVGG module, which reparameterizes the input sub-feature map and outputs it to the splicing module of the C3k2_RepVGG module; finally, the splicing module of the C3k2_RepVGG module splices and fuses the two input sub-feature maps and outputs them to the channel shuffling module for grouping, shuffling, and reorganization, and finally outputs the reorganized feature map.

4. The method for identifying key points of scallion according to claim 3, wherein: The RepVGG module adopts a multi-branch structure in the training phase and a single-branch structure in the inference phase; The multi-branch structure of the RepVGG module includes a 1x1 convolution layer, a 3x3 convolution layer, an identity mapping, and a SilU activation function. The sub-feature maps input to the RepVGG module are respectively subjected to a 1x1 convolution layer, a 3x3 convolution layer, and an identity mapping before being added and fused. The fused feature maps are then subjected to a SilU activation function, which then outputs the result to the concatenation module of the C3k2_RepVGG module. The single-branch structure adopted by the RepVGG module in the inference stage includes a 3x3 convolution layer and a SilU activation function. The sub-feature map input to the RepVGG module passes through the 3x3 convolution layer and the SilU activation function in turn and is output to the splicing module of the C3k2_RepVGG module.

5. The method for identifying key points of scallion according to claim 2, wherein: The two-layer routing attention mechanism module inputs feature maps The following processing steps are included: First, the input feature map X is divided into S×S non-overlapping regions, each of which includes feature vectors, reshape the input feature X into the input feature X r , And through linear projection, we get the query tensor Q, key tensor K, and value tensor V corresponding to the query, key, and value. The query tensor Q, key tensor K, and value tensor V are calculated based on the following formula: Q=X r W q ,K=X r W k ,V=X r W v in, is a set of real numbers, H is the height of the feature map X, W is the width of the feature map X, C is the number of channels of the feature map X, are the projection weights of query, key, and value respectively; Secondly, the query tensor Q is averaged region by region to obtain the matrix Q r , the key tensor K is averaged region by region to obtain the matrix K r , we get the matrix Q by matrix multiplication r and matrix K r The inter-region adjacency matrix of Use the directed graph method to construct inter-region routing, the inter-region adjacency matrix A r Calculated based on the following formula: A r =Q r (K r ) T Where, (K r ) T K r The transposed matrix of Again, take the first k most relevant associations in each row to prune the directed graph and construct a routing index matrix I r , routing index matrix I r Calculated based on the following formula: I r =topkIndex(A r ) In the formula, topkIndex is the adjacency matrix A between regions r Keep the associations of the top k maximum values in each row row by row; In the routing index matrix I r On the top, the token-to-token attention mechanism is applied. The token-to-token attention mechanism starts with the routing index matrix I r Collect key tensors Sum tensor Key tensor K g Sum value tensor V g Calculated based on the following formulas: K g =gather(K,I r ),V g =gather(V,I r ) Where, gather is the collection operation from the routing index matrix; The output of the token-to-token attention mechanism is calculated based on the following formula: Output=Attention(Q,K g ,V g )+LCE(V) Among them, LCE(V) is the local context enhancement term, which uses the depth convolution with a convolution kernel size of 5×5 to parameterize the value tensor V; d k K g The dimension size (K g ) T K g The transposed matrix of z is z i for The i-th element in, i ranges from 1 to L, L is length.

6. The method for identifying key points of scallion according to claim 2, wherein: The feature enhancement layer MobileNet Variants includes a depth convolution module, a first activation function layer, a first residual connection layer, a point-by-point convolution module, a second activation function layer, and a second residual connection layer connected in sequence; The first and second activation function layers are based on the following formula: Where LeakyReLU:f(x) is the first activation function layer or the second activation function, x is the feature map input to the first activation function layer or the second activation function, and α is a constant.

7. The method for identifying key points of scallion according to claim 1, wherein: The loss function Loss of the improved YOLO11 network is calculated based on the following formula: Loss=ε·CIoU Loss+β·MSE Loss Where CIoU Loss is the intersection-over-union loss function, MSE Loss is the mean square error loss function, ε is the weight coefficient of the intersection-over-union loss function, and β is the weight coefficient of the mean square error loss function.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for identifying key points of onion heads described in any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for identifying key points of onion heads described in any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for identifying key points of onion heads described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Public traffic identification method adopting improved lightweight convolutional neural network

    CN116229134A

  • Urban low-altitude small unmanned aerial vehicle detection method and system based on improved YOLOv7

    CN117115686A

  • Strawberry fruit identification method based on improved YOLOv5s

    CN118397427A

  • Method for identifying small and complex flaws of cloth, computer equipment and medium

    CN119180808A

  • Brain tumor MRI image detection method based on improved YOLOv8n

    CN119648705A