Image recognition method for automatic pepper picking machine
By introducing non-uniform scale feature extraction, color distribution dynamic mapping and cross-task feature sharing mechanisms in the image recognition method, the problems of low accuracy and poor robustness of pepper recognition in the existing technology are solved, and higher recognition accuracy and adaptability are achieved, and it is suitable for intelligent agricultural environments with resource-constrained resources.
Patent Information
- Application Number
- CN202510472968.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing image recognition methods have low recognition accuracy and poor robustness in natural environments with large lighting changes, severe fruit occlusion or complex backgrounds. The deep learning-based methods have strong sample dependence, high computing resource requirements, and difficult to guarantee real-time, which limits their actual deployment effect in resource-constrained environments.
A method of feature enhancement of pepper image based on non-uniform scale feature extraction and dynamic mapping of color distribution is proposed, and a cross-task feature sharing mechanism is added in the process of pepper image recognition to improve the information split caused by complete separation of classification and regression tasks in the existing decoupling head.
It improves the accuracy and robustness of pepper recognition, enhances the adaptability of occlusion scenes, optimizes real-time computing requirements, and provides more accurate and stable solutions for smart agriculture.
Smart Images

Figure CN120014471A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image recognition, and in particular relates to an image recognition method for an automatic pepper picking machine. Background Art
[0002] Automated harvesting technology has become an important research direction in modern facility agriculture. As a vegetable with high economic value, peppers have a high planting density and a changeable growing environment. The planting area is wide, the varieties are numerous, the fruit shapes are different, the distribution is dense, and they are easily blocked by leaves, which brings significant challenges to automated harvesting operations. The traditional manual picking method is not only labor-intensive, but also inefficient and costly, and it is difficult to meet the needs of large-scale planting for efficient and accurate harvesting. In the automatic picking system, image recognition technology, as the core of the perception link, is mainly responsible for the detection and positioning of pepper targets, which directly affects the accuracy of picking path planning and robotic arm control.
[0003] At present, mainstream image recognition methods mostly rely on traditional machine vision or target detection algorithms based on deep learning. Traditional machine vision methods are based on color thresholds, morphological features or edge information. In natural environments with large lighting changes, severe fruit occlusion or complex backgrounds, the recognition accuracy is low and the robustness is poor. Although the recognition method based on deep learning has significant advantages in target detection accuracy and adaptability, it still faces problems such as strong sample dependence, high requirements for model computing resources, and difficulty in ensuring real-time performance, which limits its actual deployment effect in resource-constrained environments. In order to solve the above problems, a feature enhancement method for pepper images based on non-uniform scale feature extraction and color distribution dynamic mapping is proposed. In the process of pepper image recognition, a cross-task feature sharing mechanism is added to improve the information fragmentation caused by the complete separation of classification and regression tasks by the existing decoupling head, improve the pepper recognition accuracy, enhance the adaptability of occluded scenes, and optimize the real-time computing requirements, providing a more accurate and stable solution for the field of smart agriculture. Summary of the invention
[0004] The present invention provides an image recognition method for an automatic pepper picking machine. Firstly, a non-uniform scale module is constructed, and a non-uniform scale transformation factor is introduced to replace the scale scaling strategy of the traditional fixed factor, so as to realize the dynamic adjustment of the scale level characteristics according to the target distribution; secondly, a scale adaptive recombination strategy is proposed to improve the traditional scale fusion method, and the dynamic scale weighting coefficient is calculated by scale response measurement to realize the weighted fusion of the scale features after channel adjustment; further, a color distribution dynamic mapping mechanism is designed, and the color characteristics of peppers at different maturity stages are effectively perceived by combining color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, so as to improve the color expression ability under complex lighting conditions; finally, the pepper image recognition module optimizes the information fragmentation problem caused by the complete separation of classification and regression tasks in the existing decoupling head through a cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition.
[0005] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: an image recognition method for an automatic pepper picking machine, comprising the following steps.
[0006] S1. Obtain images of automatic pepper pickers and create a dataset.
[0007] S2. Construct a non-uniform scaling module, improve the scaling strategy of the fixed factor through the non-uniform scaling transformation factor, and dynamically adjust the characteristics of the scale level according to the target distribution.
[0008] S3. A scale-adaptive recombination strategy is proposed to improve the traditional scale fusion method. The dynamic scale weighting coefficient is calculated according to the scale response metric to weightedly fuse the scale features after channel adjustment.
[0009] S4. Design a color distribution dynamic mapping mechanism to optimize color feature expression. Through color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, dynamically perceive the color characteristics of peppers at different stages and adapt to complex lighting environments.
[0010] S5. Construct a chili image recognition module, improve the information fragmentation caused by the existing decoupling head completely separating the classification and regression tasks through a cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer.
[0011] S6. Construct a pepper image recognition model, enhance features through non-uniform scale module, scale adaptive recombination strategy and color distribution dynamic mapping mechanism, and input the enhanced features into the pepper image recognition module to perform pepper image recognition to obtain pepper maturity and location information.
[0012] Preferably, in step S1, pepper images are collected in different environments such as sunny days, cloudy days, backlight, and backlight, with different degrees of maturity of mature peppers and immature peppers, and with different occlusion conditions such as leaf occlusion, branch occlusion, and fruit overlap. Then, pepper images with blur, severe exposure abnormality, and focus failure are eliminated, the resolution of pepper images is unified, and LabelMe is used to annotate the pepper images with bounding boxes and maturity, and the data set is divided into a training set, a validation set, and a test set.
[0013] Preferably, in step S2, the specific steps of the non-uniform scale module are as follows.
[0014] S21, firstly, dynamically adjust the distribution of scale levels by designing a scale transformation factor, wherein the scale transformation factor The specific calculation formula is: ; In the formula, is the adaptive scale offset, is the hierarchical index, , L is the maximum level index; The adaptive scale offset The specific calculation formula is: ; In the formula, is the number of small targets, is the number of large targets, is the ratio of the area of the ith target to the entire image area, N is the total number of targets, is the average size of all targets. If there are many small targets, then Become larger to retain more details. If the target is generally larger, then Become smaller so that low-level feature maps do not need to be over-scaled to reduce the amount of calculation; S22, then according to the scale transformation factor Compute nonlinear scaling function , through the nonlinear scaling function Constructing non-uniform scaling factors Optimize the scale transformation, with more flexible nonlinear characteristics, the non-uniform scale transformation factor The specific calculation formula is: ; In the formula, is a nonlinear scale mapping function, t is the training time step, and represents the current number of iterations. To control the amplitude of dynamic transformation, f is the frequency control factor; The specific calculation formula of the nonlinear scale mapping function is: ; In the formula, is the mean of all scale levels, ; S23, calculate the scaled feature size through the non-uniform scale transformation factor, and then extract features of different scales through convolution operation, and input the pepper image feature , where H, W and C are height, width and channel respectively. The height of the layer feature map He Kuan The specific calculation formula is: ; ; Use convolution operation to extract the Features of layer feature maps , and finally obtain a multi-scale feature set .
[0015] Preferably, in step S2, the non-uniform scale module adaptively adjusts the scaling factor of each scale according to the actual size of the pepper, ensuring that the model can adaptively process small targets and large targets, avoiding the limitations of traditional fixed-scale methods on small target peppers, and improving the detection accuracy of small target peppers; generating feature maps of multiple scales can ensure that the information of small target peppers is retained at a high level, and by adjusting the scale distribution, a feature map that is more suitable for small target pepper detection can be generated, thereby improving its detection accuracy under complex backgrounds, occlusion and low light conditions.
[0016] Preferably, in step S3, the specific steps of the local scale adaptive reorganization strategy are as follows.
[0017] S31. Calculate scale response metrics by local contrast of scale features Measuring the importance of scale features in multi-scale fusion, The specific calculation formula of the scale response measurement of layer scale characteristics is: ; In the formula, for Local contrast of layer-scale features, for Layer scale characteristics; The local contrast The specific calculation formula is: ; In the formula, and They are The height and width of layer-scale features, for Layer scale features The eigenvalues of the positions, for The mean of the layer scale characteristics, = ; S32, adjusting the channels of the scale features through a channel transformation matrix, wherein the channel transformation matrix adaptively adjusts the channel information of each scale through continuous integration, so that each layer of scale can effectively work together at the channel level, The layer-scale feature channel adjustment formula is: ; In the formula, After channel adjustment Layer scale characteristics, is the channel transformation matrix obtained by continuous integration; S33, calculate the dynamic scale weighting coefficient according to the scale response metric, and use the dynamic scale weighting coefficient to weight the scale features after channel adjustment to obtain the multi-scale fusion feature , , Dynamic scale weighting coefficients at layer scale The specific calculation formula is: ; In the formula, Used to normalize the scale weights to ensure that the sum of the weights of all scales is 1. is the sum of the differences between the response measure and the local contrast at all scales.
[0018] Preferably, in step S3, the scale adaptive reconstruction strategy dynamically calculates the importance of each scale feature in multi-scale fusion by introducing a scale response metric based on local contrast, so as to achieve more refined scale feature selection; combined with the channel transformation matrix in the form of continuous integral, the features of different scales are adaptively coordinated at the channel level, thereby enhancing the intrinsic correlation and complementarity between multi-scale features; finally, the dynamic scale weighting coefficient driven by the response metric and contrast is used to achieve effective weighted fusion of scale features, which not only fully retains the multi-scale detail information of small targets and complex backgrounds, but also effectively avoids the interference of redundant features, improves the feature expression ability and detection robustness, and maintains better computational efficiency, which is suitable for actual scenarios such as automatic pepper picking that have high requirements for small target detection and complex background resolution.
[0019] Preferably, in step S4, the specific steps of the color distribution dynamic mapping mechanism are as follows.
[0020] The color distribution dynamic mapping mechanism first calculates the pepper color transformation vector through adaptive transformation of color feature channels, then controls the influence of color attention through dynamic interaction modeling of color channels, and finally obtains the pepper color enhancement feature through global color adjustment. S41, color feature channel adaptive transformation adjusts the importance of color channels by introducing a dynamic color transformation matrix, the dynamic color transformation matrix The specific calculation formula is: ; In the formula, For location The interaction weights of the color channels, , , H and W are the height and width of the multi-scale fusion feature, respectively, and the interaction weight of the color channel The specific calculation formula is: ; In the formula, is a multi-scale fusion feature. It is a small convolutional network; Then use the dynamic color transformation matrix to multiply the multi-scale fusion feature to get the pepper color transformation vector ; S42. Color channel dynamic interaction modeling defines the color channel attention coefficient by the interaction relationship between color channels , the specific calculation formula is: ; The color channel attention coefficient is used to define the feature vector of the pepper after color enhancement. The calculation formula is: ; In the formula, is the feature vector of pepper after color enhancement, is element-wise multiplication, is the learnable balance factor, , is the sigmoid function, which is used to Constraints between; S43, global color adjustment, by adjusting the contribution of pepper color to the recognition of pepper maturity, obtains pepper color enhancement features , the specific calculation formula is: ; In the formula, is the color value of the pth pixel, is the global average color value, , C is the total number of pixels, are predefined reference color values.
[0021] Preferably, in step S4, the color distribution dynamic mapping mechanism first designs an adaptive transformation of color feature channels, dynamically adjusts color channels according to pixel positions and color characteristics, and adapts color components to different task requirements; secondly, introduces color channel dynamic interaction modeling, and dynamically models the coupling relationship of color channels, so that color information can be more effectively integrated; finally, global color adjustment obtains pepper color enhancement features by adjusting the contribution of pepper color to pepper maturity recognition; improves robustness to challenging scenes such as different maturity, occlusion, and lighting changes; compared with traditional color processing methods, this method can adaptively adjust the importance of color features instead of fixedly using RGB color channels, thereby improving detection accuracy and reducing false detections.
[0022] Preferably, in step S5, the pepper image recognition module has the following specific steps.
[0023] S51, input color enhancement feature , the cross-task feature sharing mechanism enhances the color feature Classification Features and regression features , define the cross-task sharing coefficient according to the uncertainty measurement of the classification task and the regression task, and calculate the shared features through the cross-task sharing coefficient , the specific calculation formula is: ; In the formula, is the sigmoid function, is the cross-task sharing coefficient, is the feature conversion matrix from classification to regression, is the feature conversion matrix from regression to classification; The cross-task sharing coefficient calculation formula is: ; In the formula, is the uncertainty measure for the classification task, is the uncertainty measure for the regression task, is the sensitivity adjustment factor. When the classification task fluctuates greatly, , becomes smaller, resulting in Tends to 1, enhancing the feature sharing of classification tasks. When the regression task fluctuates greatly, , becomes smaller, resulting in Tends to 0, enhancing the feature sharing of regression tasks. When the difficulty of two tasks is similar, , feature balanced sharing; S52, the classification branch assigns a category label to the peppers in the pepper image by combining the classification features and the shared features to determine whether the peppers are ripe; the regression branch predicts the position of the target bounding box using the regression features and the shared features to determine the precise position of the peppers in the image; the pepper image recognition is optimized by designing a joint optimization loss function by adding the classification-regression consistency loss and the pepper maturity loss on the basis of the traditional loss function; S53. Design a joint optimization loss function to ensure the consistency of the detection frame and the segmentation mask, optimize the pepper image recognition effect, and the specific calculation formula of the joint optimization loss function is: ; In the formula, is the weight of the traditional classification and regression loss function, is the weight of the classification-regression consistency loss function, is the weight of the pepper maturity loss function, is the traditional classification and regression loss function, is the classification-regression consistency loss function, is the pepper maturity loss function; S54. Calculate the classification-regression consistency loss function by the difference between the detection box and the segmentation mask. The specific calculation formula is: ; In the formula, is the predicted bounding box of the jth sample, is the segmentation mask of the jth sample, is the true bounding box of the jth sample, is the overlap ratio between the detection box and the segmentation mask, is the overlap ratio between the ground-truth bounding box and the segmentation mask; S55, calculating the pepper maturity loss function by analyzing the pepper color, the specific calculation formula is: ; In the formula, is the predicted value of maturity of the mth sample, is the true maturity label of the mth sample, 0 represents immature and 1 represents mature. is the color feature of the mth sample, is the maturity threshold.
[0024] Preferably, in step S5, the classification task and the regression task share some underlying features through shared feature learning, thereby improving information complementarity and enabling the detection and segmentation tasks to enhance each other; the classification-regression consistency loss function ensures the consistency of the detection box and the segmentation mask, and improves the overall robustness of the model. At the same time, the pepper maturity loss function makes the picking decision more intelligent and reduces the mispicking rate and missed picking rate; the joint optimization loss function can improve the pepper recognition accuracy, enhance the adaptability to occluded scenes, and optimize the real-time computing requirements.
[0025] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides an image recognition method for an automatic pepper picking machine. A non-uniform scale module realizes dynamic adjustment of scale level features according to target distribution by introducing non-uniform scale transformation factors to replace the traditional fixed factor scale scaling strategy; a scale adaptive recombination strategy is proposed to improve the traditional scale fusion method, and a dynamic scale weighting coefficient is calculated by scale response measurement to realize weighted fusion of scale features after channel adjustment; a color distribution dynamic mapping mechanism is further designed, and color characteristics of peppers at different maturity stages are effectively perceived by combining color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, so as to improve the color expression ability under complex lighting conditions; finally, the pepper image recognition module optimizes the information fragmentation problem caused by the complete separation of classification and regression tasks in the existing decoupling head through a cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 The present invention provides a flowchart of an image recognition method for an automatic pepper picking machine.
[0027] Figure 2 It is a structural diagram of the scale-adaptive reorganization strategy provided by the present invention.
[0028] Figure 3 It is a structural diagram of a pepper image recognition module provided by the present invention.
[0029] Figure 4 This is a pepper image recognition effect diagram provided by the present invention. DETAILED DESCRIPTION
[0030] The following will be combined with the accompanying drawings of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] See also Figures 1 to 4 The present invention provides an image recognition method for an automatic pepper picking machine, constructs a non-uniform scale module, introduces a non-uniform scale transformation factor to replace the traditional fixed factor scale scaling strategy, and realizes the dynamic adjustment of the scale level characteristics according to the target distribution; proposes a scale adaptive recombination strategy to improve the traditional scale fusion method, calculates the dynamic scale weighting coefficient through scale response measurement, and realizes the weighted fusion of scale features after channel adjustment; further designs a color distribution dynamic mapping mechanism, combines color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, effectively perceives the color characteristics of peppers at different maturity stages, and improves the color expression ability under complex lighting conditions; finally, the pepper image recognition module optimizes the information fragmentation problem caused by the complete separation of classification and regression tasks in the existing decoupling head through a cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition.
[0032] See also Figure 1 As shown, an image recognition method for an automatic pepper picking machine in an embodiment of the present application.
[0033] S1. Obtain images of automatic pepper pickers and create a dataset.
[0034] Furthermore, pepper images were collected in different environments such as sunny days, cloudy days, backlight, and backlight, and in different maturity levels of mature peppers and immature peppers, and in different occlusion conditions such as leaf occlusion, branch occlusion, and fruit overlap. Then, the pepper images with blur, severe exposure abnormalities, and focus failures were removed, and 5,000 normal pepper images were obtained, including 2,000 mature peppers, 2,000 immature peppers, 500 occluded pepper images, and 500 images under special lighting conditions. The pepper image resolution was unified to 640×640px, and LabelMe was used to annotate the pepper images with bounding boxes and maturity levels. The dataset was divided into training set, validation set, and test set according to the ratio of 7:1.5:1.5.
[0035] S2. Construct a non-uniform scaling module, improve the scaling strategy of the fixed factor through the non-uniform scaling transformation factor, and dynamically adjust the characteristics of the scale level according to the target distribution.
[0036] Furthermore, the specific steps of the non-uniform scale module are as follows.
[0037] S21, firstly, dynamically adjust the distribution of scale levels by designing a scale transformation factor, wherein the scale transformation factor The specific calculation formula is: ; In the formula, is the adaptive scale offset, is the hierarchical index, , L is the maximum level index, set to 8; The adaptive scale offset The specific calculation formula is: ; In the formula, is the number of small targets, the initial value is set to 1, is the number of large targets, the initial value is set to 1, is the ratio of the area of the ith target to the entire image area, N is the total number of targets, and the initial value is set to 2. is the average size of all targets. If there are many small targets, then Become larger to retain more details. If the target is generally larger, then Become smaller so that low-level feature maps do not need to be over-scaled to reduce the amount of computation.
[0038] S22, then according to the scale transformation factor Compute nonlinear scaling function , through the nonlinear scaling function Constructing non-uniform scaling factors Optimize the scale transformation, with more flexible nonlinear characteristics, the non-uniform scale transformation factor The specific calculation formula is: ; In the formula, is a nonlinear scale mapping function, t is the training time step, represents the current number of iterations, and the total number of iterations is 50. To control the amplitude of dynamic transformation, set the range to , the initial value is set to 0.1, f is the frequency control factor, and the setting range is , the initial value is set to ; The specific calculation formula of the nonlinear scale mapping function is: ; In the formula, is the mean of all scale levels, .
[0039] S23, calculate the scaled feature size through the non-uniform scale transformation factor, and then extract features of different scales through convolution operation, and input the pepper image feature , where H, W and C are height, width and channel respectively, height is 640, width is 640, channel is 3, for the The height of the layer feature map He Kuan The specific calculation formula is: , ; , ; Use convolution operation to extract the Features of layer feature maps , and finally obtain a multi-scale feature set .
[0040] S3. A scale-adaptive recombination strategy is proposed to improve the traditional scale fusion method. The dynamic scale weighting coefficient is calculated according to the scale response metric to weightedly fuse the scale features after channel adjustment.
[0041] Furthermore, if Figure 2 As shown, the specific steps of the scale-adaptive reconstruction strategy are as follows.
[0042] S31. Calculate scale response metrics by local contrast of scale features Measuring the importance of scale features in multi-scale fusion, The specific calculation formula of the scale response measurement of layer scale characteristics is: ; In the formula, for Local contrast of layer-scale features, for Layer scale characteristics; The local contrast The specific calculation formula is: ; In the formula, and They are The height and width of layer-scale features, for Layer scale features The eigenvalues of the positions, for The mean of the layer scale characteristics, = .
[0043] S32, adjusting the channels of the scale features through a channel transformation matrix, wherein the channel transformation matrix adaptively adjusts the channel information of each scale through continuous integration, so that each layer of scale can effectively work together at the channel level, The layer-scale feature channel adjustment formula is: ; In the formula, After channel adjustment Layer scale characteristics, is the channel transformation matrix obtained by continuous integration.
[0044] S33, calculate the dynamic scale weighting coefficient according to the scale response metric, and use the dynamic scale weighting coefficient to weight the scale features after channel adjustment to obtain the multi-scale fusion feature , , Dynamic scale weighting coefficients at layer scale The specific calculation formula is: ; In the formula, Used to normalize the scale weights to ensure that the sum of the weights of all scales is 1. is the sum of the differences between the response measure and the local contrast at all scales.
[0045] S4. Design a color distribution dynamic mapping mechanism to optimize color feature expression. Through color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, dynamically perceive the color characteristics of peppers at different stages and adapt to complex lighting environments.
[0046] Furthermore, the specific steps of the color distribution dynamic mapping mechanism are as follows.
[0047] The color distribution dynamic mapping mechanism first calculates the pepper color transformation vector through adaptive transformation of color feature channels, then controls the influence of color attention through dynamic interaction modeling of color channels, and finally obtains the pepper color enhancement feature through global color adjustment.
[0048] S41, color feature channel adaptive transformation adjusts the importance of color channels by introducing a dynamic color transformation matrix, the dynamic color transformation matrix The specific calculation formula is: ; In the formula, For location The interaction weights of the color channels, , , the interaction weights of the color channels The specific calculation formula is: ; In the formula, is a multi-scale fusion feature. It is a small convolutional network, using 3×3 convolution and 1×1 convolution; Then use the dynamic color transformation matrix to multiply the multi-scale fusion feature to get the pepper color transformation vector .
[0049] S42. Color channel dynamic interaction modeling defines the color channel attention coefficient by the interaction relationship between color channels , the specific calculation formula is: ; The color channel attention coefficient is used to define the feature vector of the pepper after color enhancement. The calculation formula is: ; In the formula, is the feature vector of pepper after color enhancement, is element-wise multiplication, is the learnable balance factor, , is the sigmoid function, which is used to Constraints between.
[0050] S43, global color adjustment, by adjusting the contribution of pepper color to the recognition of pepper maturity, obtains pepper color enhancement features , the specific calculation formula is: ; In the formula, is the color value of the pth pixel, is the global average color value, , C is the total number of pixels, set to 640×640, is a predefined reference color value, .
[0051] S5. Construct a chili image recognition module, improve the information fragmentation caused by the existing decoupling head completely separating the classification and regression tasks through a cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer.
[0052] Furthermore, the specific steps of the pepper image recognition module are as follows.
[0053] S51, input color enhancement feature , the cross-task feature sharing mechanism enhances the color feature Classification Features and regression features , define the cross-task sharing coefficient according to the uncertainty measurement of the classification task and the regression task, and calculate the shared features through the cross-task sharing coefficient , the specific calculation formula is: ; In the formula, is the sigmoid function, is the cross-task sharing coefficient, is the feature conversion matrix from classification to regression, is the feature conversion matrix from regression to classification; The cross-task sharing coefficient calculation formula is: ; In the formula, is the uncertainty measure for the classification task, , For a batch of samples The amount of change, , is the uncertainty measure for the regression task, , M is the number of samples in the training batch, set to 3500, is the gradient value of the regression boundary, is the sensitivity adjustment factor. When the classification task fluctuates greatly, , becomes smaller, resulting in Tends to 1, enhancing the feature sharing of classification tasks. When the regression task fluctuates greatly, , becomes smaller, resulting in Tends to 0, enhancing the feature sharing of regression tasks. When the difficulty of two tasks is similar, , feature balanced sharing.
[0054] S52. The classification branch assigns category labels to the peppers in the pepper image by combining classification features and shared features to determine whether the peppers are ripe. The regression branch uses regression features and shared features to predict the position of the target bounding box and determine the exact position of the peppers in the image. The pepper image recognition is optimized by designing a joint optimization loss function by adding classification-regression consistency loss and pepper maturity loss to the traditional loss function.
[0055] S53. Design a joint optimization loss function to ensure the consistency of the detection frame and the segmentation mask, optimize the pepper image recognition effect, and the specific calculation formula of the joint optimization loss function is: ; In the formula, is the weight of the traditional classification and regression loss function, is the weight of the classification-regression consistency loss function, is the weight of the pepper maturity loss function, is the traditional classification and regression loss function, is the classification-regression consistency loss function, is the pepper maturity loss function.
[0056] S54. Calculate the classification-regression consistency loss function by the difference between the detection box and the segmentation mask. The specific calculation formula is: ; In the formula, is the predicted bounding box of the jth sample, is the segmentation mask of the jth sample, is the true bounding box of the jth sample, is the overlap ratio between the detection box and the segmentation mask, is the overlap ratio between the ground-truth bounding box and the segmentation mask.
[0057] S55, calculating the pepper maturity loss function by analyzing the pepper color, the specific calculation formula is: ; In the formula, is the predicted value of maturity of the mth sample, is the true maturity label of the mth sample, 0 represents immature and 1 represents mature. is the color feature of the mth sample, is the maturity threshold, set to .
[0058] S6. Construct a pepper image recognition model, enhance features through non-uniform scale module, scale adaptive recombination strategy and color distribution dynamic mapping mechanism, and input the enhanced features into the pepper image recognition module to perform pepper image recognition to obtain pepper maturity and location information.
[0059] Furthermore, if Figure 3 As shown, in step S6, the pepper image is input into the pepper image recognition model, and the features are first divided into 8 layers of features through the non-uniform scale module, and the multi-scale features are fused using the scale adaptive recombination strategy. The fused features are enhanced in color to achieve feature enhancement of the pepper image, and then the enhanced features are input into the pepper image recognition module for recognition, and the pepper image recognition result is output; the pepper image recognition model is based on the Pytorch framework and is implemented through the Pycharm application. The model is trained using 3500 training images in the pepper image dataset, and the trained model is tested using the test set and verified on the verification set. The pepper image recognition model improves the pepper image recognition accuracy while also improving the recognition efficiency, making the pepper image recognition model more suitable for real-time pepper picking.
[0060] Furthermore, if Figure 4 As shown in the figure, the recognition effect of the pepper image recognition model is shown, including a ripe pepper and an unripe pepper, with the pepper position border and maturity information marked.
[0061] The above are only preferred embodiments of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. An image recognition method for an automatic pepper picking machine, characterized in that: The following steps are involved: S1, obtain images of automatic pepper picking machines and create a data set; S2, build a non-uniform scale module, improve the scale scaling strategy of the fixed factor through the non-uniform scale transformation factor, and dynamically adjust the characteristics of the scale level according to the target distribution; S3. A scale adaptive recombination strategy is proposed to improve the traditional scale fusion method. The scale features after channel adjustment are weighted and fused by calculating the dynamic scale weighting coefficient according to the scale response metric. S4. Design a color distribution dynamic mapping mechanism to optimize color feature expression, dynamically perceive the color characteristics of peppers at different stages and adapt to complex lighting environments through color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment; S5. Construct a pepper image recognition module, improve the information fragmentation caused by the existing decoupling head completely separating the classification and regression tasks through the cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer; S6. Construct a pepper image recognition model, enhance features through non-uniform scale module, scale adaptive recombination strategy and color distribution dynamic mapping mechanism, and input the enhanced features into the pepper image recognition module to perform pepper image recognition to obtain pepper maturity and location information.
2. The image recognition method for an automatic pepper picking machine according to claim 1, characterized in that: In the step S1, images of the automatic pepper picking machine are obtained and made into a data set, pepper images in different environments, different maturity, and different occlusion conditions are collected, and the pepper images with blur, serious exposure abnormality, and focus failure are eliminated. Then, the pepper image resolution is unified, and the pepper images are annotated with bounding boxes and maturity using LabelMe, and the data set is divided into a training set, a validation set, and a test set.
3. The image recognition method for an automatic pepper picking machine according to claim 2, characterized in that: In the step S2, the specific steps of the non-uniform scale module are: S21, firstly, dynamically adjust the distribution of scale levels by designing a scale transformation factor, wherein the scale transformation factor The specific calculation formula is: ; In the formula, is the adaptive scale offset, is the hierarchical index, , L is the maximum level index; The adaptive scale offset The specific calculation formula is: ; In the formula, is the number of small targets, is the number of large targets, is the ratio of the area of the i-th target to the area of the entire pepper image, and N is the total number of targets; S22, then according to the scale transformation factor Compute nonlinear scaling function , through the nonlinear scaling function Constructing non-uniform scaling factors Optimize the scale transformation, the non-uniform scale transformation factor The specific calculation formula is: ; In the formula, is a nonlinear scale mapping function, t is the training time step, To control the amplitude of dynamic transformation, f is the frequency control factor; The specific calculation formula of the nonlinear scale mapping function is: ; In the formula, is the mean of all scale levels; S23, calculate the scaled feature size through the non-uniform scale transformation factor, and then extract features of different scales through convolution operation, and finally obtain a multi-scale feature set .
4. The image recognition method for an automatic pepper picking machine according to claim 3 is characterized in that: In the step S3, the specific steps of the scale-adaptive reorganization strategy are as follows: S31. Calculate scale response metrics by local contrast of scale features Measuring the importance of scale features in multi-scale fusion, The specific calculation formula of the scale response measurement of layer scale characteristics is: ; In the formula, for Local contrast of layer-scale features, for Layer scale characteristics; The local contrast The specific calculation formula is: ; In the formula, and They are The height and width of layer-scale features, for Layer scale features The eigenvalues of the positions, for Mean of layer-scale characteristics; S32, adjusting the channels of the scale feature through a channel transformation matrix, wherein the channel transformation matrix adaptively adjusts the channel information of each scale through a continuous integration method, The layer-scale feature channel adjustment formula is: ; In the formula, After channel adjustment Layer scale characteristics; S33, calculate the dynamic scale weighting coefficient according to the scale response metric, and use the dynamic scale weighting coefficient to weight the scale features after channel adjustment to obtain the multi-scale fusion feature , , Dynamic scale weighting coefficients at layer scale The specific calculation formula is: 。 5. The image recognition method for an automatic pepper picking machine according to claim 4, characterized in that: In the step S4, the specific steps of the color distribution dynamic mapping mechanism are as follows: The color distribution dynamic mapping mechanism first calculates the pepper color transformation vector through adaptive transformation of color feature channels, then controls the influence of color attention through dynamic interaction modeling of color channels, and finally obtains the pepper color enhancement feature through global color adjustment. S41, color feature channel adaptive transformation adjusts the importance of color channels by introducing a dynamic color transformation matrix, which is composed of the interaction weights of color channels Then use the dynamic color transformation matrix to multiply the multi-scale fusion feature to get the pepper color transformation vector , the interaction weights of the color channels The specific calculation formula is: ; In the formula, is a multi-scale fusion feature. is a small convolutional network. , , H and W are the height and width of the multi-scale fusion features respectively; S42. Color channel dynamic interaction modeling defines the color channel attention coefficient by the interaction relationship between color channels , the specific calculation formula is: ; The color channel attention coefficient is used to define the feature vector of the pepper after color enhancement. The calculation formula is: ; In the formula, is the feature vector of pepper after color enhancement, is element-wise multiplication, is the learnable balance factor; S43. Global color adjustment: By adjusting the contribution of pepper color to the recognition of pepper maturity, the pepper color enhancement feature is obtained. , the specific calculation formula is: ; In the formula, is the color value of the pth pixel, is the global average color value, C is the total number of pixels, are predefined reference color values.
6. The image recognition method for an automatic pepper picking machine according to claim 5, characterized in that: In the step S5, the pepper image recognition module specifically comprises the following steps: S51, input color enhancement feature , the cross-task feature sharing mechanism enhances the color feature Classification features and regression features , define the cross-task sharing coefficient according to the uncertainty measurement of the classification task and the regression task, and calculate the shared features through the cross-task sharing coefficient , the specific calculation formula is: ; In the formula, is the sigmoid function, is the cross-task sharing coefficient, is the feature conversion matrix from classification to regression, is the feature conversion matrix from regression to classification; The cross-task sharing coefficient calculation formula is: ; In the formula, is the uncertainty measure for the classification task, is the uncertainty measure for the regression task, is the sensitivity adjustment factor; S52. The classification branch assigns category labels to the peppers in the pepper image by combining classification features and shared features to determine whether the peppers are ripe. The regression branch uses regression features and shared features to predict the position of the target bounding box and determine the exact position of the peppers in the image. The pepper image recognition is optimized by designing a joint optimization loss function by adding classification-regression consistency loss and pepper maturity loss to the traditional loss function.
Citation Information
Patent Citations
Fruit and vegetable picking robot and target identification picking method thereof
CN119054511A
Lightweight real-time infrared target detection method and system for unmanned ship
CN119600513A
Method for recognizing distribution network equipment based on raspberry pi multi-scale feature fusion
US11631238B1