An image recognition method for a chili automatic picking machine

By adopting non-uniform scale feature extraction, color distribution dynamic mapping and cross-task feature sharing mechanism image recognition methods in the pepper automatic picking system, the problems of low accuracy and poor robustness of pepper recognition in the existing technology are solved, and higher recognition accuracy and adaptability are achieved, and real-time computing needs are optimized.

CN120014471BActive Publication Date: 2025-07-01GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510472968.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-01
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing image recognition methods have low accuracy and poor robustness in the pepper automatic picking system, especially in environments where light changes are large, fruit occlusion is severe or background complex, and deep learning-based methods have problems such as strong sample dependence, high computing resources, and difficult to guarantee real-time.

Method used

A method of image feature enhancement based on non-uniform scale feature extraction and color distribution dynamic mapping is proposed. Combined with the cross-task feature sharing mechanism, it improves the information splitting problem of classification and regression tasks in existing decoupling heads.

Benefits of technology

It improves the accuracy and robustness of pepper recognition, enhances the adaptability of occlusion scenarios, optimizes real-time computing requirements, and provides more accurate and stable solutions for smart agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014471B_ABST
    Figure CN120014471B_ABST
Patent Text Reader

Abstract

The present invention provides an image recognition method for a pepper automatic picking machine, which relates to the field of image recognition. It introduces a non-uniform scale transformation factor to replace the traditional fixed-factor scale scaling strategy, realizing the dynamic adjustment of the scale level features according to the target distribution; improves the traditional scale fusion method, calculates the dynamic scale weighting coefficient through scale response measurement, and realizes the weighted fusion of the scale features after channel adjustment; combines the adaptive transformation of the color feature channel, the dynamic interaction modeling of the color channels and the global color adjustment, effectively perceives the color features of peppers at different maturity stages, and improves the color expression ability under complex lighting conditions; optimizes the information fragmentation problem caused by the complete separation of the classification and regression tasks in the existing decoupling head through a cross-task feature sharing mechanism, introduces a cross-task sharing coefficient to enhance the effective feature interaction between the classification and regression tasks in the deep feature layer, and improves the accuracy and robustness of recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and particularly relates to an image recognition method for a pepper automatic picking machine. Background Art

[0002] Automated picking technology has become an important research direction in modern facility agriculture. As a vegetable with relatively high economic value, peppers are densely planted and have a variable growth environment. They have a wide planting area, a large variety of cultivars, different fruit shapes, are densely distributed, and are easily blocked by leaves, which poses significant challenges to automated picking operations. The traditional manual picking method not only has a high labor intensity, but also has low efficiency and high costs, making it difficult to meet the requirements of large-scale planting for efficient and precise harvesting. In an automatic picking system, as the core of the perception link, image recognition technology is mainly responsible for the detection and positioning of pepper targets, directly affecting the accuracy of picking path planning and robotic arm control.

[0003] Currently, the mainstream image recognition methods mostly rely on traditional machine vision or object detection algorithms based on deep learning. Traditional machine vision methods based on color thresholds, morphological features, or edge information have low recognition accuracy and poor robustness in natural environments with large lighting changes, severe fruit occlusion, or complex backgrounds. Although the recognition methods based on deep learning have significant advantages in terms of object detection accuracy and adaptability, they still face problems such as strong sample dependence, high requirements for model computing resources, and difficulty in ensuring real-time performance, which limit their actual deployment effects in resource-constrained environments. To solve the above problems, a pepper image feature enhancement method based on non-uniform scale feature extraction and dynamic mapping of color distribution is proposed, and a cross-task feature sharing mechanism is added during the pepper image recognition process to improve the information fragmentation caused by the complete separation of the existing decoupled head for classification and regression tasks, improve the pepper recognition accuracy, enhance the adaptability to occlusion scenarios, and optimize the real-time computing requirements, providing a more accurate and stable solution for the field of intelligent agriculture. Summary of the Invention

[0004] The present invention provides an image recognition method for a chili automatic picking machine. First, a non-uniform scale module is constructed, and a non-uniform scale transformation factor is introduced to replace the scale scaling strategy of the traditional fixed factor, so as to dynamically adjust the features of the scale level according to the target distribution. Secondly, a scale adaptive recombination strategy is proposed to improve the traditional scale fusion method. By calculating the dynamic scale weighting coefficient through the scale response metric, the weighted fusion of the scale features after channel adjustment is realized. Furthermore, a color distribution dynamic mapping mechanism is designed, which combines the adaptive transformation of the color feature channel, the dynamic interaction modeling of the color channel and the global color adjustment to effectively perceive the color features of chili peppers at different maturity stages and improve the color expression ability under complex lighting conditions. Finally, the chili image recognition module optimizes the information fragmentation problem caused by the complete separation of the classification and regression tasks in the existing decoupled head through the cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between the classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition.

[0005] To achieve the above object, the present invention provides the following technical solutions: An image recognition method for a chili automatic picking machine, comprising the following steps.

[0006] S1. Obtain the images of the chili automatic picking machine and make them into a data set.

[0007] S2. Construct a non-uniform scale module, improve the scale scaling strategy of the fixed factor through the non-uniform scale transformation factor, and dynamically adjust the features of the scale level according to the target distribution.

[0008] S3. Propose a scale adaptive recombination strategy to improve the traditional scale fusion method, and calculate the dynamic scale weighting coefficient according to the scale response metric to weight and fuse the scale features after channel adjustment.

[0009] S4. Design a color distribution dynamic mapping mechanism to optimize the color feature expression, and dynamically perceive the color features of chili peppers at different stages and adapt to complex lighting environments through the adaptive transformation of the color feature channel, the dynamic interaction modeling of the color channel and the global color adjustment.

[0010] S5. Construct a chili image recognition module, improve the information fragmentation caused by the complete separation of the classification and regression tasks in the existing decoupled head through the cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer.

[0011] S6. Construct a chili image recognition model, enhance the features through the non-uniform scale module, the scale adaptive recombination strategy and the color distribution dynamic mapping mechanism, and input the enhanced features into the chili image recognition module to perform chili image recognition to obtain the chili maturity and position information.

[0012] Preferably, in step S1, pepper images are collected under different environments such as sunny days, cloudy days, backlight, and backlighting, with different maturity levels of mature peppers and immature peppers, and different occlusion conditions such as leaf occlusion, branch occlusion, and fruit overlap. Then, pepper images that are blurred, severely abnormally exposed, or have failed focus are removed, the resolution of the pepper images is unified, and LabelMe is used to perform bounding box annotation and maturity annotation on the pepper images. The dataset is divided into a training set, a validation set, and a test set.

[0013] Preferably, in step S2, the specific steps of the non-uniform scale module are as follows.

[0014] S21. First, the distribution of scale levels is dynamically adjusted by designing a scale transformation factor. The scale transformation factor The specific calculation formula is:

[0015] ;

[0016] In the formula, is the adaptive scale offset, is the level index, , L is the maximum level index;

[0017] The adaptive scale offset The specific calculation formula is:

[0018] ;

[0019] In the formula, is the number of small targets, is the number of large targets, is the proportion of the area of the i-th target to the area of the entire image, N is the total number of targets, is the average scale of all targets. If the number of small targets is large, then becomes larger to retain more details. If the targets are generally large, then becomes smaller so that the low-level feature map does not need to be overly scaled to reduce the computational amount;

[0020] S22. Then, according to the scale transformation factor calculate the non-linear scale mapping function , and construct a non-uniform scale transformation factor through the non-linear scale mapping function to optimize the scale transformation, which has more flexible non-linear characteristics. The non-uniform scale transformation factor The specific calculation formula is:

[0021] ;

[0022] In the formula, is a non - linear scale mapping function, t is the training time step, representing the current number of iterations, is the amplitude for controlling the dynamic transformation, and f is the frequency control factor;

[0023] The specific calculation formula of the non - linear scale mapping function is:

[0024] ;

[0025] In the formula, is the mean of all scale levels, ;

[0026] S23. Calculate the scaled feature size through the non - uniform scale transformation factor, then extract features of different scales through convolution operations, and input the pepper image features , where H, W, and C are the height, width, and channels respectively. For the height and width of the feature map of the

[0027] layer, the specific calculation formula is:

[0028] ;

[0029] Use convolution operations to extract the features of the layer feature map , and finally obtain the multi - scale feature set .

[0030] Preferably, in step S2, the non - uniform scale module adaptively adjusts the scaling factor of each scale according to the actual size of the pepper, ensuring that the model can adaptively process small and large targets, avoiding the limitation of the traditional fixed - scale method on small - target peppers, and improving the detection accuracy of small - target peppers; generating feature maps of multiple scales can ensure that the information of small - target peppers is retained at high levels. By adjusting the scale distribution, feature maps more suitable for small - target pepper detection can be generated, improving its detection accuracy under complex backgrounds, occlusions, and low - light conditions.

[0031] Preferably, in step S3, the specific steps of the local - scale adaptive recombination strategy are as follows.

[0032] S31. Calculate the scale response metric through the local contrast of the scale features to measure the importance of the scale features in multi - scale fusion. The scale response metric of the scale features of the

[0033] layer is calculated as follows:

[0034] In the formula, For the local contrast of the layer-scale features, is the layer-scale features;

[0035] The said local contrast has the specific calculation formula as follows:

[0036] ;

[0037] In the formula, and are respectively the height and width of the layer-scale features, is the feature value of the layer-scale features at the position, is the mean value of the layer-scale features, = ;

[0038] S32. Adjust the channels of the scale features through the channel transformation matrix. The channel transformation matrix adaptively adjusts the channel information of each scale through continuous integration, enabling each layer scale to work effectively in cooperation at the channel level. The layer-scale feature channel adjustment formula is:

[0039] ;

[0040] In the formula, is the layer-scale feature after channel adjustment, is the channel transformation matrix obtained through continuous integration;

[0041] S33. Calculate the dynamic scale weighting coefficient according to the scale response metric, and weight and fuse the scale features after channel adjustment through the dynamic scale weighting coefficient to obtain the multi-scale fusion feature , . The dynamic scale weighting coefficient of the layer scale has the specific calculation formula as follows:

[0042] ;

[0043] In the formula, is used to normalize the scale weights to ensure that the sum of the weights of all scales is 1, is the sum of the differences between the response metrics of all scales and the local contrast.

[0044] Preferably, in step S3, the scale adaptive recombination strategy can achieve more refined scale feature selection by introducing a scale response metric based on local contrast, dynamically calculating the importance of each scale feature in multi-scale fusion, and combining a continuous integral form of channel transformation matrix to enable the features of different scales to adaptively cooperate at the channel level, enhancing the internal correlation and complementarity between multi-scale features. Finally, through the dynamic scale weighting coefficient jointly driven by the response metric and contrast, the effective weighted fusion of scale features is realized, not only fully retaining the multi-scale detail information under small targets and complex backgrounds, but also effectively avoiding the interference of redundant features, improving the feature expression ability and detection robustness, while maintaining relatively good computational efficiency, and is applicable to the actual scenario of automatic pepper picking with high requirements for small target detection and complex background resolution.

[0045] Preferably, in step S4, the specific steps of the color distribution dynamic mapping mechanism are as follows.

[0046] The color distribution dynamic mapping mechanism first calculates the pepper color transformation vector through the adaptive transformation of the color feature channel, then controls the influence on color attention through the dynamic interaction modeling of the color channel, and finally obtains the enhanced pepper color feature through global color adjustment.

[0047] S41. The adaptive transformation of the color feature channel adjusts the importance of the color channel by introducing a dynamic color transformation matrix, and the dynamic color transformation matrix The specific calculation formula is:

[0048] ;

[0049] In the formula, is the interaction weight of the color channel at position , , H and W are respectively the height and width of the multi-scale fusion feature, and the interaction weight of the color channel has the specific calculation formula:

[0050] ;

[0051] In the formula, is the multi-scale fusion feature, is the small convolutional network;

[0052] Then, the dynamic color transformation matrix is multiplied by the multi-scale fusion feature to obtain the pepper color transformation vector ;

[0053] S42. The dynamic interaction modeling of the color channel defines the color channel attention coefficient through the interaction relationship between color channels, and the specific calculation formula is:

[0054] ;

[0055] Define the pepper feature vector after color enhancement through the color channel attention coefficient. The calculation formula is:

[0056] ;

[0057] In the formula, is the pepper feature vector after color enhancement, is the element-wise multiplication, is the learnable balance factor, , is the sigmoid function, which is used to constrain within ;

[0058] S43. Global color adjustment obtains the enhanced pepper color feature by adjusting the contribution of pepper color to the recognition of pepper maturity , and the specific calculation formula is:

[0059] ;

[0060] In the formula, is the color value of the p-th pixel, is the global average color value, , C is the total number of pixels, is the predefined reference color value.

[0061] Preferably, in step S4, the color distribution dynamic mapping mechanism first designs an adaptive transformation of the color feature channel, dynamically adjusts the color channel according to the pixel position and color characteristics, so that the color components adapt to different task requirements; secondly, introduces a dynamic interaction modeling of the color channel, dynamically models the coupling relationship of the color channel, so that the color information can be more effectively fused; finally, global color adjustment obtains the enhanced pepper color feature by adjusting the contribution of pepper color to the recognition of pepper maturity; improves the robustness to challenging scenarios such as different maturities, occlusions, and illumination changes; compared with traditional color processing methods, this method can adaptively adjust the importance of color features instead of fixedly using the RGB color channel, thereby improving the detection accuracy and reducing false detections.

[0062] Preferably, in step S5, the specific steps of the pepper image recognition module are as follows.

[0063] S51. Input the enhanced color feature , and the cross-task feature sharing mechanism divides the enhanced color feature into the classification feature and the regression feature , define the cross - task sharing coefficient according to the uncertainty measures of the classification task and the regression task, and calculate the shared features through the cross - task sharing coefficient , and the specific calculation formula is:

[0064] ;

[0065] In the formula, is the sigmoid function, is the cross - task sharing coefficient, is the feature transformation matrix from classification to regression, is the feature transformation matrix from regression to classification;

[0066] The calculation formula of the cross - task sharing coefficient is:

[0067] ;

[0068] In the formula, is the uncertainty measure of the classification task, is the uncertainty measure of the regression task, is the sensitivity adjustment factor. When the classification task fluctuates greatly, , becomes smaller, resulting in tending to 1, enhancing the feature sharing of the classification task. When the regression task fluctuates greatly, , becomes smaller, resulting in tending to 0, enhancing the feature sharing of the regression task. When the two tasks are of comparable difficulty, , the features are evenly shared;

[0069] S52. The classification branch assigns category labels to the peppers in the pepper image by combining the classification features and the shared features to determine whether the peppers are ripe; the regression branch uses the regression features and the shared features to predict the position of the target bounding box to determine the exact position of the peppers in the image; design a joint optimization loss function to optimize the pepper image recognition by adding the classification - regression consistency loss and the pepper ripeness loss on the basis of the traditional loss function;

[0070] S53. Design a joint optimization loss function to ensure the consistency of the detection box and the segmentation mask and optimize the pepper image recognition effect. The specific calculation formula of the joint optimization loss function is:

[0071] ;

[0072] In the formula, is the weight of the traditional classification and regression loss functions, is the weight of the classification - regression consistency loss function, is the weight of the pepper ripeness loss function, is the traditional classification and regression loss function, is the classification-regression consistency loss function, is the chili maturity loss function;

[0073] S54. Calculate the classification-regression consistency loss function by detecting the difference between the detection box and the segmentation mask. The specific calculation formula is:

[0074] ;

[0075] In the formula, is the predicted bounding box of the j-th sample, is the segmentation mask of the j-th sample, is the ground truth bounding box of the j-th sample, is the ratio of the overlap between the detection box and the segmentation mask, is the ratio of the overlap between the ground truth bounding box and the segmentation mask;

[0076] S55. Calculate the chili maturity loss function by analyzing the color of the chili. The specific calculation formula is:

[0077] ;

[0078] In the formula, is the predicted maturity value of the m-th sample, is the ground truth maturity label of the m-th sample, where 0 represents immature and 1 represents mature, is the color feature of the m-th sample, is the maturity threshold.

[0079] Preferably, in step S5, through shared feature learning, the classification task and the regression task share a part of the underlying features, improving information complementarity, so that the detection and segmentation tasks can enhance each other; the classification-regression consistency loss function ensures the consistency between the detection box and the segmentation mask, improving the overall robustness of the model. At the same time, the chili maturity loss function makes the picking decision more intelligent, reducing the mis-picking rate and the missed-picking rate; the joint optimization loss function can improve the chili recognition accuracy, enhance the adaptability to occlusion scenarios, and optimize the real-time calculation requirements.

[0080] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0081] The present invention provides an image recognition method for a chili automatic picking machine. The non-uniform scale module realizes the feature of dynamically adjusting the scale level according to the target distribution by introducing a non-uniform scale transformation factor to replace the scale scaling strategy of the traditional fixed factor; a scale adaptive recombination strategy is proposed to improve the traditional scale fusion method, and the dynamic scale weighting coefficient is calculated through the scale response metric to realize the weighted fusion of the scale features after channel adjustment; further, a color distribution dynamic mapping mechanism is designed, which combines the adaptive transformation of the color feature channel, the dynamic interaction modeling of the color channel and the global color adjustment to effectively perceive the color features of chili peppers at different maturity stages and improve the color expression ability under complex lighting conditions; finally, the chili image recognition module optimizes the information fragmentation problem caused by the complete separation of the classification and regression tasks in the existing decoupled head through the cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between the classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 FIG. is a flowchart of an image recognition method for a chili automatic picking machine provided by the present invention.

[0083] Figure 2 FIG. is a structural diagram of a scale adaptive recombination strategy provided by the present invention.

[0084] Figure 3 FIG. is a structural diagram of a chili image recognition module provided by the present invention.

[0085] Figure 4 FIG. is an effect diagram of chili image recognition provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0087] Please refer to Figures 1 to 4The present invention provides an image recognition method for an automatic pepper picking machine, constructs a non-uniform scale module, introduces a non-uniform scale transformation factor to replace the traditional fixed factor scale scaling strategy, and realizes the dynamic adjustment of the scale level characteristics according to the target distribution; proposes a scale adaptive recombination strategy to improve the traditional scale fusion method, calculates the dynamic scale weighting coefficient through scale response measurement, and realizes the weighted fusion of scale features after channel adjustment; further designs a color distribution dynamic mapping mechanism, combines color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment, effectively perceives the color characteristics of peppers at different maturity stages, and improves the color expression ability under complex lighting conditions; finally, the pepper image recognition module optimizes the information fragmentation problem caused by the complete separation of classification and regression tasks in the existing decoupling head through a cross-task feature sharing mechanism, and introduces a cross-task sharing coefficient to enhance the effective feature interaction between classification and regression tasks in the deep feature layer, thereby improving the accuracy and robustness of recognition.

[0088] See also Figure 1 As shown, an image recognition method for an automatic pepper picking machine in an embodiment of the present application.

[0089] S1. Obtain images of automatic pepper pickers and create a dataset.

[0090] Furthermore, pepper images were collected in different environments such as sunny days, cloudy days, backlight, and backlight, and in different maturity levels of mature peppers and immature peppers, and in different occlusion conditions such as leaf occlusion, branch occlusion, and fruit overlap. Then, the pepper images with blur, severe exposure abnormalities, and focus failures were removed, and 5,000 normal pepper images were obtained, including 2,000 mature peppers, 2,000 immature peppers, 500 occluded pepper images, and 500 images under special lighting conditions. The pepper image resolution was unified to 640×640px, and LabelMe was used to annotate the pepper images with bounding boxes and maturity levels. The dataset was divided into training set, validation set, and test set according to the ratio of 7:1.5:1.5.

[0091] S2. Construct a non-uniform scaling module, improve the scaling strategy of the fixed factor through the non-uniform scaling transformation factor, and dynamically adjust the characteristics of the scale level according to the target distribution.

[0092] Furthermore, the specific steps of the non-uniform scale module are as follows.

[0093] S21, firstly, dynamically adjust the distribution of scale levels by designing a scale transformation factor, wherein the scale transformation factor The specific calculation formula is:

[0094] ;

[0095] In the formula, is the adaptive scale offset, is the hierarchical index, , where L is the maximum hierarchical index, set to 8;

[0096] The adaptive scale offset The specific calculation formula is:

[0097] ;

[0098] In the formula, is the number of small targets, with the initial value set to 1, is the number of large targets, with the initial value set to 1, is the ratio of the area of the i-th target to the area of the entire image, N is the total number of targets, with the initial value set to 2, is the average scale of all targets. If the number of small targets is large, then becomes larger to retain more details. If the targets are generally large, then becomes smaller so that the low-level feature map does not need to be scaled excessively to reduce the computational load.

[0099] S22. Then, according to the scale transformation factor calculate the non-linear scale mapping function . Through the non-linear scale mapping function construct the non-uniform scale transformation factor to optimize the scale transformation, which has more flexible non-linear characteristics. The non-uniform scale transformation factor The specific calculation formula is:

[0100] ;

[0101] In the formula, is the non-linear scale mapping function, t is the training time step, representing the current iteration number. The total number of iterations is 50, is the amplitude controlling the dynamic transformation, with the setting range of , and the initial value is set to 0.1. f is the frequency control factor, with the setting range of , and the initial value is set to ;

[0102] The specific calculation formula of the non-linear scale mapping function is:

[0103] ;

[0104] In the formula, is the mean of all scale levels, .

[0105] S23. Calculate the scaled feature size through a non-uniform scale transformation factor, then extract features of different scales through a convolution operation, and input the features of the chili image. , where H, W, and C are the height, width, and number of channels respectively, the height is 640, the width is 640, and the number of channels is 3. For the height and width of the

[0106] -th layer feature map, the specific calculation formula is:

[0107] , ;

[0108] Use a convolution operation to extract the features of the -th layer feature map, and finally obtain a multi-scale feature set .

[0109] S3. Propose a scale-adaptive recombination strategy to improve the traditional scale fusion method, and calculate the dynamic scale weighting coefficient according to the scale response metric to weight and fuse the scale features after channel adjustment.

[0110] Furthermore, as Figure 2 shown, the specific steps of the scale-adaptive recombination strategy are as follows.

[0111] S31. Calculate the scale response metric through the local contrast of the scale features to measure the importance of the scale features in multi-scale fusion. The specific calculation formula of the scale response metric of the -th

[0112] layer scale features is:

[0113] In the formula, is the local contrast of the -th layer scale features, is the

[0114] -th layer scale features;

[0115] The specific calculation formula of the local contrast

[0116] is: and are the height and width of the -th layer scale features respectively, is the feature value of the -th is the mean of the layer-scale features, = .

[0117] S32. Adjust the channels of the scale features through a channel transformation matrix. The channel transformation matrix adaptively adjusts the channel information of each scale through continuous integration, enabling each layer scale to work effectively together at the channel level. The formula for adjusting the channels of the layer-scale features is:

[0118] ;

[0119] In the formula, is the layer-scale feature after channel adjustment, and is the channel transformation matrix obtained through continuous integration.

[0120] S33. Calculate the dynamic scale weighting coefficient according to the scale response metric, and fuse the scale features after channel adjustment through the dynamic scale weighting coefficient to obtain the multi-scale fusion feature , , and the dynamic scale weighting coefficient of the layer scale is specifically calculated as:

[0121] ;

[0122] In the formula, is used to normalize the scale weights to ensure that the sum of the weights of all scales is 1, is the sum of the differences between the response metrics of all scales and the local contrast.

[0123] S4. Design a dynamic color distribution mapping mechanism to optimize the color feature expression, and dynamically perceive the color features of peppers at different stages and adapt to complex lighting environments through color feature channel adaptive transformation, color channel dynamic interaction modeling, and global color adjustment.

[0124] Furthermore, the specific steps of the dynamic color distribution mapping mechanism are as follows.

[0125] The dynamic color distribution mapping mechanism first calculates the pepper color transformation vector through color feature channel adaptive transformation, then controls the influence on color attention through color channel dynamic interaction modeling, and finally obtains the pepper color enhancement feature through global color adjustment.

[0126] S41. The color feature channel adaptive transformation adjusts the importance of the color channels by introducing a dynamic color transformation matrix. The dynamic color transformation matrix is specifically calculated as:

[0127] ;

[0128] Wherein, is the interaction weight of the position color channel, , , the interaction weight of the said color channel The specific calculation formula is:

[0129] ;

[0130] Wherein, is the multi-scale fusion feature, is a small convolutional network, using 3×3 convolution and 1×1 convolution;

[0131] Then multiply the dynamic color transformation matrix with the multi-scale fusion feature to obtain the pepper color transformation vector .

[0132] S42. Color channel dynamic interaction modeling defines the color channel attention coefficient through the interaction relationship between color channels, and the specific calculation formula is:

[0133] ;

[0134] Define the pepper feature vector after color enhancement through the color channel attention coefficient, and the calculation formula is:

[0135] ;

[0136] Wherein, is the pepper feature vector after color enhancement, is element-wise multiplication, is a learnable balance factor, , is the sigmoid function, used to constrain within between.

[0137] S43. Global color adjustment obtains the pepper color enhancement feature by adjusting the contribution of pepper color to pepper maturity recognition, and the specific calculation formula is:

[0138] ;

[0139] Wherein, is the color value of the p-th pixel, is the global average color value, , C is the total number of pixels, set to 640×640, is a predefined reference color value, .

[0140] S5. Build a chili image recognition module, improve the information fragmentation caused by the complete separation of classification and regression tasks in the existing decoupled head through the cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer.

[0141] Furthermore, the specific steps of the chili image recognition module are as follows.

[0142] S51. Input color-enhanced features , and the cross-task feature sharing mechanism divides the color-enhanced features into classification features and regression features . Define the cross-task sharing coefficient according to the uncertainty measures of the classification task and the regression task, and calculate the shared features through the cross-task sharing coefficient . The specific calculation formula is:[[]]

[0143] ;

[0144] In the formula, is the sigmoid function, is the cross-task sharing coefficient, is the feature transformation matrix from classification to regression, is the feature transformation matrix from regression to classification;

[0145] The calculation formula of the cross-task sharing coefficient is:[[]]

[0146] ;

[0147] In the formula, is the uncertainty measure of the classification task, , is the variation of a batch of samples, , is the uncertainty measure of the regression task, , M is the number of samples in the training batch, set to 3500, is the gradient value of the regression boundary, is the sensitivity adjustment factor. When the classification task fluctuates greatly, , becomes smaller, resulting in tending to 1 and enhancing the feature sharing of the classification task. When the regression task fluctuates greatly, , becomes smaller, resulting in tending to 0 and enhancing the feature sharing of the regression task. When the difficulties of the two tasks are comparable, , feature equilibrium sharing.

[0148] S52. The classification branch assigns a category label to the chili pepper in the chili pepper image by combining the classification feature and the shared feature, and determines whether the chili pepper is ripe; the regression branch uses the regression feature and the shared feature to predict the position of the target bounding box to determine the exact position of the chili pepper in the image; the joint optimization loss function is designed by adding the classification-regression consistency loss and the chili pepper ripeness loss on the basis of the traditional loss function to optimize the chili pepper image recognition.

[0149] S53. Design the joint optimization loss function to ensure the consistency of the detection box and the segmentation mask, and optimize the chili pepper image recognition effect. The specific calculation formula of the joint optimization loss function is:

[0150] ;

[0151] In the formula, is the weight of the traditional classification and regression loss function, is the weight of the classification-regression consistency loss function, is the weight of the chili pepper ripeness loss function, is the traditional classification and regression loss function, is the classification-regression consistency loss function, is the chili pepper ripeness loss function.

[0152] S54. Calculate the classification-regression consistency loss function through the difference between the detection box and the segmentation mask. The specific calculation formula is:

[0153] ;

[0154] In the formula, is the predicted bounding box of the j-th sample, is the segmentation mask of the j-th sample, is the true bounding box of the j-th sample, is the ratio of the overlap between the detection box and the segmentation mask, is the ratio of the overlap between the true bounding box and the segmentation mask.

[0155] S55. Calculate the chili pepper ripeness loss function by analyzing the color of the chili pepper. The specific calculation formula is:

[0156] ;

[0157] In the formula, is the predicted ripeness value of the m-th sample, is the true ripeness label of the m-th sample, 0 means unripe, 1 means ripe, is the color feature of the m-th sample, is the maturity threshold, set to .

[0158] S6. Construct a pepper image recognition model, enhance features through a non-uniform scale module, a scale-adaptive recombination strategy, and a color distribution dynamic mapping mechanism, and input the enhanced features into the pepper image recognition module to perform pepper image recognition to obtain pepper maturity and position information.

[0159] Furthermore, as Figure 3 shown, in step S6, the pepper image is input into the pepper image recognition model. First, the features are divided into 8 layers of features through the non-uniform scale module, and the multi-scale features are fused using the scale-adaptive recombination strategy. The fused features are subjected to color feature enhancement to achieve feature enhancement of the pepper image. Then, the enhanced features are input into the pepper image recognition module for recognition, and the pepper image recognition result is output. The pepper image recognition model is based on the Pytorch framework and is implemented through the Pycharm application program. The model is trained using 3500 training images in the pepper image dataset. The trained model is tested using the test set and verified on the validation set. The pepper image recognition model improves both the accuracy and efficiency of pepper image recognition, making the pepper image recognition model more suitable for real-time pepper picking.

[0160] Furthermore, as Figure 4 shown, the figure shows the recognition effect of the pepper image recognition model, including a mature pepper and an immature pepper, with the pepper position border and maturity information marked.

[0161] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. An image recognition method for an automatic pepper picking machine, characterized in that: The following steps are involved: S1, obtain images of automatic pepper picking machines and create a data set; S2. Construct a non-uniform scale module. Improve the scale scaling strategy of the fixed factor through the non-uniform scale transformation factor. Dynamically adjust the characteristics of the scale level according to the target distribution. The specific steps of constructing the non-uniform scale module are as follows: S21, first, dynamically adjust the distribution of scale levels by designing a scale transformation factor, wherein the scale transformation factor λ l The specific calculation formula is: In the formula, δ l is the adaptive scale offset, l is the level index, l∈{1, 2, ..., L}, L is the maximum level index; The adaptive scale offset δ l The specific calculation formula is: Where N small is the number of small targets, N large is the number of large targets, A i is the ratio of the area of ​​the i-th target to the area of ​​the entire pepper image, and N is the total number of targets; S22, then according to the scale transformation factor λ l Calculate the nonlinear scaling function λ′ l , through the nonlinear scaling function λ′ l Constructing non-uniform scaling factors Optimize the scale transformation, the non-uniform scale transformation factor The specific calculation formula is: In the formula, λ′ l is the nonlinear scale mapping function, t is the training time step, γ l To control the amplitude of dynamic transformation, f is the frequency control factor; The specific calculation formula of the nonlinear scale mapping function is: In the formula, λ μ is the mean of all scale levels; S23, calculate the scaled feature size through the non-uniform scale transformation factor, and then extract features of different scales through convolution operation, and finally obtain the multi-scale feature set F mult ={F′1,F′2,...,F′ L }; S3. A scale adaptive recombination strategy is proposed to improve the traditional scale fusion method. The scale features after channel adjustment are weighted and fused according to the scale response metric by calculating the dynamic scale weighting coefficient. The scale response metric ρ l To measure the importance of scale features in multi-scale fusion, the specific calculation formula of the scale response metric of the l-layer scale feature is: ρ l =∫0 1 (δ(F′ l )) 0.3 dx; In the formula, δ(F′ l ) is the local contrast of the l-level scale feature, F l ′ is the scale feature of level l; S4. Design a color distribution dynamic mapping mechanism to optimize color feature expression, dynamically perceive the color characteristics of peppers at different stages and adapt to complex lighting environments through color feature channel adaptive transformation, color channel dynamic interaction modeling and global color adjustment; S5. Construct a pepper image recognition module, improve the information fragmentation caused by the existing decoupling head completely separating the classification and regression tasks through the cross-task feature sharing mechanism, and design a cross-task sharing coefficient to optimize the feature interaction between classification and regression in the deep feature layer; S6. Construct a pepper image recognition model, enhance features through non-uniform scale module, scale adaptive recombination strategy and color distribution dynamic mapping mechanism, and input the enhanced features into the pepper image recognition module to perform pepper image recognition to obtain pepper maturity and location information.

2. The image recognition method for an automatic pepper picking machine according to claim 1, characterized in that: In the step S1, images of the automatic pepper picking machine are obtained and made into a data set, pepper images in different environments, different maturity, and different occlusion conditions are collected, and the pepper images with blur, serious exposure abnormality, and focus failure are eliminated. Then, the pepper image resolution is unified, and the pepper images are annotated with bounding boxes and maturity using LabelMe, and the data set is divided into a training set, a validation set, and a test set.

3. The image recognition method for an automatic pepper picking machine according to claim 2, characterized in that: In the step S3, the specific steps of the scale-adaptive reorganization strategy are as follows: S31. Calculate the scale response metric ρ by the local contrast of the scale feature l , the local contrast δ(F l The specific calculation formula of ′) is: In the formula, H l and W l are the height and width of the l-th layer scale feature, F l ′(h, w) is the eigenvalue of the l-layer scale feature at the position (h, w), ε l is the mean of the scale feature of layer l; S32, adjusting the channels of the scale feature through the channel transformation matrix, the channel transformation matrix adaptively adjusts the channel information of each scale through continuous integration, and the l-layer scale feature channel adjustment formula is: In the formula, is the l-layer scale feature after channel adjustment; S33, calculate the dynamic scale weighting coefficient according to the scale response metric, and use the dynamic scale weighting coefficient to weight the scale features after channel adjustment to obtain the multi-scale fusion feature H and W are the height and width of the multi-scale fusion feature, respectively, C is the total number of pixels, and the dynamic scale weighting coefficient ω of the l-level scale l The specific calculation formula is:

4. The image recognition method for an automatic pepper picking machine according to claim 3 is characterized in that: In the step S4, the specific steps of the color distribution dynamic mapping mechanism are as follows: The color distribution dynamic mapping mechanism first calculates the pepper color transformation vector through adaptive transformation of color feature channels, then controls the influence of color attention through dynamic interaction modeling of color channels, and finally obtains the pepper color enhancement feature through global color adjustment. S41, color feature channel adaptive transformation adjusts the importance of color channels by introducing a dynamic color transformation matrix, which is composed of the interaction weights ω of the color channels ij (x, y), and then the dynamic color transformation matrix is ​​multiplied by the multi-scale fusion feature to obtain the pepper color transformation vector C trans , the interaction weights ω of the color channels ij (x, y) The specific calculation formula is: ij (x, y) = tanh(f θ (F all , x, y)); In the formula, F all is the multi-scale fusion feature, f θ is a small convolutional network, x∈[0,H], y∈[0,W], H and W are the height and width of the multi-scale fusion features respectively; S42. Color channel dynamic interaction modeling defines the color channel attention coefficient A through the interaction relationship between color channels. color , the specific calculation formula is: The color channel attention coefficient is used to define the feature vector of the pepper after color enhancement. The calculation formula is: C enhanced =A color ⊙C trans +b trans ·C trans ; In the formula, C enhanced is the feature vector of pepper after color enhancement, ⊙ is the element-by-element multiplication, β trans is the learnable balance factor; S43, global color adjustment By adjusting the contribution of pepper color to the recognition of pepper maturity, the pepper color enhancement feature C is obtained. final , the specific calculation formula is: In the formula, C p is the color value of the pth pixel, C ave is the global average color value, C is the total number of pixels, and C ref are predefined reference color values.

5. The image recognition method for an automatic pepper picking machine according to claim 4, characterized in that: In the step S5, the pepper image recognition module specifically comprises the following steps: S51. Input color enhancement feature C final , the cross-task feature sharing mechanism enhances the color feature C final Divided into classification features F D and regression feature F S , the cross-task sharing coefficient is defined according to the uncertainty measurement of the classification task and the regression task, and the shared feature F is calculated by the cross-task sharing coefficient shard , the specific calculation formula is: F shard =σ(α shard ·W DS ·F D +(1-a shard )·W SD ·F S ); In the formula, σ is the sigmoid function, α shard is the cross-task sharing coefficient, W DS is the feature conversion matrix from classification to regression, W SD is the feature conversion matrix from regression to classification; The cross-task sharing coefficient calculation formula is: Where D det is the uncertainty measure for the classification task, D seg is the uncertainty measure of the regression task, λ s is the sensitivity adjustment factor; S52. The classification branch assigns category labels to the peppers in the pepper image by combining classification features and shared features to determine whether the peppers are ripe. The regression branch uses regression features and shared features to predict the position of the target bounding box and determine the exact position of the peppers in the image. The pepper image recognition is optimized by designing a joint optimization loss function by adding classification-regression consistency loss and pepper maturity loss to the traditional loss function.

Citation Information

Patent Citations

  • Fruit and vegetable picking robot and target identification picking method thereof

    CN119054511A

  • Lightweight real-time infrared target detection method and system for unmanned ship

    CN119600513A