Apple identification method in complex environment

Through the deep neural network method of image enhancement, attention mechanism and multi-scale feature fusion, the problems of branch and leaf occlusion, scale diversity and lighting changes of apple recognition in complex orchard environments are solved, and the accurate identification of apple location and category is achieved, and the intelligence level of orchard management is improved.

CN120279546AInactive Publication Date: 2025-07-08新疆理工学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357295.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex orchard environments, the apple identification task faces problems such as branch and leaf occlusion, serious background interference, apple scale diversity and lighting changes, which make it difficult for traditional detection methods to accurately locate targets and are prone to missed or missed detection.

Method used

Image enhancement technology is used to preprocess the original image, and the preset annotation tool is used to accurately mark Apple location and category, and an attention mechanism and multi-scale feature fusion method are introduced in deep neural networks. Model training is combined with improved loss functions to suppress background interference and deal with Apples of different scales.

Benefits of technology

It significantly improves the accuracy and robustness of apple recognition in orchard environment, realizes accurate identification of apple location and category, and supports intelligent orchard management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279546A_ABST
    Figure CN120279546A_ABST
Patent Text Reader

Abstract

The invention discloses an apple identification method in a complex environment, and the method comprises the steps: obtaining an original image containing an apple, and carrying out the preprocessing, and obtaining a first image; marking the position and category of the apple in the first image to generate a second image; inputting the second image into a deep neural network, introducing an attention mechanism, calculating the attention weight of the apple region, and generating a third feature map; based on the third feature map, adopting a multi-scale feature fusion method, and combining high-level semantic information and low-level space details to generate a fourth feature map; reconstructing a network neck module, integrating feature information of different scales, and generating a fifth feature map; and performing model training based on the fifth feature map, deploying the trained model to an orchard environment, identifying the position and category of the apple in real time, and outputting a result. According to the method, the accuracy and robustness of apple recognition in the orchard environment are remarkably improved, and powerful support is provided for intelligent orchard management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of agricultural intelligent detection, and particularly relates to an apple recognition method in a complex environment. Background Art

[0002] In a complex orchard environment, the apple recognition task faces many technical problems. First of all, apples are usually blocked by branches and leaves in the image, and the background interference is serious. The branches and leaves in the orchard are lush, the colors of apples and branches and leaves are similar, and the pixels of the robot camera are low. These factors result in blurred and dark images, missing edge information of apple contours, and seriously affecting the prediction accuracy of the detection network. In this case, traditional detection methods are difficult to accurately locate the target, and it is easy to have missed detections or false detections.

[0003] Secondly, apples present variable morphological and color characteristics under different lighting, weather and growth stages. Unstable lighting conditions, such as backlighting, shadows or uneven lighting, will cause changes in the color and shape characteristics of apples. In addition, there are differences in the size, shape and color of apples at different growth stages. For example, young fruits are smaller and lighter in color, while mature fruits are larger and darker in color. These variable characteristics increase the difficulty of recognition, making a single detection model difficult to adapt to all scenarios.

[0004] In addition, the scale of apples in the image varies greatly, especially the detection accuracy of small target apples is low. In a complex orchard environment, the sizes of apples are different, and they may present different scales due to the distance from the camera. For small target apples, the recall rate of traditional detection models is low, and it is easy to have missed detections. At the same time, when the apples are similar in color to the background or blocked, the detection difficulty is further increased.

[0005] To solve these problems, it is urgent to propose an apple recognition method in a complex environment. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes an apple recognition method in a complex environment to solve the problems existing in the above prior art.

[0007] To achieve the above object, the present invention provides an apple recognition method in a complex environment, including the following steps:

[0008] Obtain the original image containing apples, and preprocess the original image by using image enhancement technology to obtain the first enhanced image;

[0009] According to the apple regions in the first image, use a preset annotation tool to annotate the position and category of each apple, and generate a second image containing the target position and category information;

[0010] Input the second image into the deep neural network, introduce an attention mechanism module in the feature extraction module, calculate the attention weights of the apple regions, and obtain a third feature map with attention weights;

[0011] Based on the third feature map, adopt a multi-scale feature fusion method in the feature extraction module, combine high-level semantic information and low-level spatial details, and generate a fourth feature map containing multi-scale information;

[0012] According to the fourth feature map, reconstruct the neck module of the deep neural network, adopt a feature fusion mechanism, integrate feature information of different scales, and obtain a fifth feature map after fusion;

[0013] Train the deep neural network based on the fused fifth feature map. During the model training process, adopt an improved loss function, combine the problems of target boundary blur and scale change, calculate the loss value between the prediction result and the true annotation, and obtain an optimized loss value;

[0014] Through the backpropagation algorithm, transfer the optimized loss value into the deep neural network, update the model parameters, and obtain a trained apple recognition model;

[0015] Deploy the trained apple recognition model to the orchard environment, obtain orchard images in real time, input them into the completed apple recognition model, judge the position and category of apples in the images, and output the recognition results.

[0016] Optionally, for obtaining the original image containing apples and preprocessing the original image using an image enhancement technique to obtain an enhanced first image, it includes:

[0017] Obtain the original image containing apples in the orchard environment, and for the foliage occlusion and background interference in the original image, adopt the histogram equalization algorithm to enhance the image and obtain an enhanced first image.

[0018] Optionally, for accurately annotating the position and category of each apple according to the apple regions in the first image using a preset annotation tool to generate a second image containing target position and category information, it includes:

[0019] Adopt preset rules to divide the region of the first image to determine the distribution range of apple targets in the image;

[0020] According to the target position information, call the annotation tool to draw bounding boxes for the apple regions;

[0021] If the target category information matches the preset category, perform category marking on the pixel points within the annotation box;

[0022] Optimize the boundaries of the annotation boxes through a preset threshold, obtain the information of the optimized annotation boxes, and then generate a second image containing the target location and category information.

[0023] Optionally, input the second image into a deep neural network, introduce an attention mechanism module in the feature extraction module, and obtain a third feature map with attention weights by calculating the attention weights of the apple region, including:

[0024] Use a deep neural network to extract features from the second image, embed an attention mechanism module in the feature extraction module; calculate the attention weight values of the apple region, and at the same time suppress the interference of the background domain and the foliage domain to obtain a third feature map with attention weights.

[0025] Optionally, based on the third feature map, adopt a multi-scale feature fusion method in the feature extraction module, combine high-level semantic information and low-level spatial details, and generate a fourth feature map containing multi-scale information, including:

[0026] Adopt a multi-scale feature fusion method in the feature extraction module to extract features from the input third feature map, and obtain high-level semantic features and low-level spatial features;

[0027] For the high-level semantic features and low-level spatial features, use a feature fusion module for weighted fusion to generate a fourth feature map containing multi-scale information.

[0028] Optionally, according to the fourth feature map, reconstruct the neck module of the neural network, adopt a feature fusion mechanism, integrate feature information of different scales, and obtain a fused fifth feature map, including:

[0029] Optimize the fourth feature map using a reconstruction method and adjust the structure of the neural network neck module;

[0030] For the multi-scale feature information in the fourth feature map, design a feature fusion mechanism to achieve the integration of different-scale features; through the information integration process, obtain a fused fifth feature map.

[0031] Optionally, during the model training process, adopt an improved loss function, combine the problems of target boundary blur and scale change, calculate the loss value between the prediction result and the true annotation, and obtain an optimized loss value, including:

[0032] Analyze the characteristics of target boundary blur and scale change, and extract the blur degree of the target boundary and the key points of scale change;

[0033] Combine the target boundary information to correct the deviation of the prediction result and adjust the influence of scale change, and recalculate the loss value;

[0034] Adjust the weight of the loss function according to the correlation between the target boundary blur and scale change, and optimize the calculation process of the loss value;

[0035] Dynamically adjust the parameters of the loss function by analyzing the blur problem and extracting the characteristics of the change problem;

[0036] Until the maximum number of iterations is reached, obtain the optimized loss value.

[0037] The present invention also provides an apple recognition system in a complex environment for implementing the above method, and the system includes:

[0038] An image preprocessing module for obtaining an original image containing apples in an orchard environment, and preprocessing the original image using an image enhancement technique to obtain a first enhanced image;

[0039] A labeling generation module for accurately labeling the position and category of each apple according to the apple area in the first image using a preset labeling tool, and generating a second image containing target position and category information;

[0040] A feature extraction module for inputting the second image into a deep neural network, introducing an attention mechanism module in the feature extraction module, and obtaining a third feature map with attention weights by calculating the attention weights of the apple area;

[0041] A multi-scale feature fusion module for generating a fourth feature map containing multi-scale information based on the third feature map, adopting a multi-scale feature fusion method in the feature extraction module, and combining high-level semantic information and low-level spatial details;

[0042] A feature reconstruction module for reconstructing the neck module of the deep neural network according to the fourth feature map, adopting a feature fusion mechanism to integrate feature information of different scales, and obtaining a fifth fused feature map;

[0043] A loss calculation module for training the deep neural network based on the fused fifth feature map. During the model training process, an improved loss function is adopted, combined with the problems of target boundary blur and scale change, to calculate the loss value between the prediction result and the true annotation, and obtain the optimized loss value;

[0044] A model training module for transmitting the optimized loss value to the neural network through the backpropagation algorithm, updating the model parameters, and obtaining a trained apple recognition model;

[0045] An identification output module for deploying the trained apple recognition model to the orchard environment, real-time acquiring orchard images, inputting them into the model, judging the position and category of apples in the images, and outputting the recognition result.

[0046] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method.

[0047] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.

[0048] Compared with the prior art, the present invention has the following advantages and technical effects:

[0049] The present invention discloses an apple recognition method in a complex environment. Aiming at problems such as foliage occlusion, background interference, and apple scale diversity in orchards, the present invention uses image enhancement technology to preprocess the original image, and uses a preset annotation tool to accurately annotate the position and category of apples. An attention mechanism and a multi-scale feature fusion method are introduced into the deep neural network to effectively suppress background interference and process apples of different scales. By reconstructing the neck module of the neural network, an efficient feature fusion mechanism is adopted to integrate multi-scale information. During the training process, an improved loss function is used to deal with the problems of fuzzy object boundaries and scale changes. Finally, the trained model is deployed to the orchard environment to achieve accurate recognition of the position and category of apples. The present invention significantly improves the accuracy and robustness of apple recognition in the orchard environment and provides strong support for intelligent orchard management. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0051] Figure 1 is a schematic flowchart of the apple recognition method according to an embodiment of the present invention;

[0052] Figure 2 is a schematic diagram of the second image generation process according to an embodiment of the present invention;

[0053] Figure 3 is a schematic flowchart of the backpropagation process according to an embodiment of the present invention;

[0054] Figure 4 is a schematic diagram of the system structure according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0056] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0057] Embodiment 1

[0058] As Figure 1 shown, in this embodiment, an apple recognition method in a complex environment is provided, including the following steps:

[0059] Obtain the original image containing apples, and preprocess the original image using image enhancement technology to obtain the enhanced first image;

[0060] According to the apple regions in the first image, use a preset annotation tool to annotate the positions and categories of each apple, and generate a second image containing target position and category information;

[0061] Input the second image into a deep neural network, introduce an attention mechanism module in the feature extraction module, calculate the attention weights of the apple regions, and obtain a third feature map with attention weights;

[0062] Based on the third feature map, adopt a multi-scale feature fusion method in the feature extraction module, combine high-level semantic information and low-level spatial details, and generate a fourth feature map containing multi-scale information;

[0063] According to the fourth feature map, reconstruct the neck module of the deep neural network, adopt a feature fusion mechanism, integrate feature information of different scales, and obtain a fused fifth feature map;

[0064] Based on the fused fifth feature map, train the deep neural network. During the model training process, adopt an improved loss function, combine the problems of target boundary blur and scale change, calculate the loss value between the prediction result and the true annotation, and obtain an optimized loss value;

[0065] Through the backpropagation algorithm, transfer the optimized loss value into the deep neural network, update the model parameters, and obtain a trained apple recognition model;

[0066] Deploy the trained apple recognition model to the orchard environment, obtain orchard images in real time, input them into the completed apple recognition model, judge the positions and categories of apples in the images, and output the recognition results.

[0067] Feasible. The acquisition of the original image containing apples uses image enhancement technology to preprocess the original image to obtain the enhanced first image, including: acquiring the original image containing apples in the orchard environment, and for the foliage occlusion and background interference in the original image, using the histogram equalization algorithm to enhance the image to obtain the enhanced first image.

[0068] As a specific embodiment, the original image acquired in the orchard may be affected by uneven light, foliage occlusion, etc. The histogram equalization can effectively improve the image quality. For example, a photo of an apple tree taken at dusk has an overall dark image and insufficient contrast due to insufficient light. After histogram equalization, the gray levels of the image can be remapped to make the gray distribution more uniform, thereby making the targets such as apples and foliage clearer.

[0069] Feasible, such as Figure 2 shown, according to the apple regions in the first image, using a preset annotation tool to accurately annotate the position and category of each apple to generate a second image containing target position and category information, including:

[0070] Using preset rules to divide the region of the first image to determine the distribution range of apple targets in the image; according to the target position information, calling the annotation tool to draw a bounding box for the apple region; if the target category information matches the preset category, then perform category marking on the pixel points within the annotation box; optimizing the boundary of the annotation box through a preset threshold to obtain the optimized annotation box information, and then generating a second image containing target position and category information.

[0071] As a specific implementation manner, in the image region division, the preset rules are usually set based on the color features and spatial distribution characteristics of the image. For example, when apples are red or yellow, the hue range can be set between 0 degrees and 60 degrees, the saturation threshold is above 50%, and the brightness threshold is between 30% and 90%. Such preset rules can effectively identify the apple target regions in the fruit trees. The call of the annotation tool needs to consider the morphological characteristics of different fruits. For round apples, a rectangular bounding box can be used for annotation, and the size of the box is automatically adjusted according to the target size. When encountering an occluded apple, the bounding box needs to contain the maximum visible range to ensure the integrity of the annotation.

[0072] During the category marking process, the system matches the predefined category information with the recognition results. For example, ripe apples are marked as the first category, and unripe apples are marked as the second category. Inside the annotation box, each pixel will be assigned a corresponding category attribute value for subsequent object tracking and recognition. The optimization of the annotation box involves improving the boundary accuracy. By setting the boundary threshold, such as setting the boundary tolerance to plus or minus five pixels, the system can automatically adjust the position and size of the annotation box. This optimization can effectively eliminate the errors in manual annotation and improve the accuracy of object positioning.

[0073] During the generation process of the second image, the system superimposes the optimized annotation information on the original image. Each annotation box contains the position coordinates and category number of the object. For example, the upper left corner coordinates of the annotation box of an apple are 250 pixels and 300 pixels, the width of the box is 80 pixels, the height is 90 pixels, and the category number is 1.

[0074] Morphological filtering can smooth the edges of the annotation box. By setting the size of the structural element, such as a 3x3 rectangular kernel, dilation and erosion operations are performed on the annotation box. This can eliminate the jagged effect of the boundary and make the annotation box more regular.

[0075] Implementable, inputting the second image into a deep neural network, introducing an attention mechanism module in the feature extraction module, and obtaining a third feature map with attention weights by calculating the attention weights of the apple region, including:

[0076] Using a deep neural network to extract features from the second image, embedding an attention mechanism module in the feature extraction module; calculating the attention weight values of the apple region, and at the same time suppressing the interference of the background domain and the foliage domain to obtain a third feature map with attention weights.

[0077] As a specific implementation, when the deep neural network extracts features from the second image, a multi-layer convolutional structure is adopted, and different sizes of convolutional kernels are set in each layer to extract features of different scales. For example, a 3x3 convolutional kernel is used in the first layer to extract edge features, a 5x5 convolutional kernel is used in the second layer to extract texture features, and a 7x7 convolutional kernel is used in the third layer to extract shape features.

[0078] Introduce an attention mechanism during the feature extraction process, and obtain the attention weights by calculating the similarity matrix of each region of the feature map. For the calculation of the attention weights of the apple region, a combination of spatial attention and channel attention is adopted. Spatial attention focuses on the position information of the apple, and channel attention focuses on the color and texture features of the apple.

[0079] For the interference suppression of the background and branches and leaves, an adaptive threshold suppression mechanism is introduced. The initial threshold is set to 0.6. When the attention weight of the background area is detected to be greater than this threshold, its weight value is attenuated to below 0.3. For the branch and leaf area, morphological opening operation is used to remove small connecting structures, and then the weight value of this area is reduced to below 0.2. This can effectively highlight the feature expression of the apple area.

[0080] Implementable, based on the third feature map, a multi-scale feature fusion method is adopted in the feature extraction module to combine high-level semantic information and low-level spatial details to generate a fourth feature map containing multi-scale information, including:

[0081] In the feature extraction module, a multi-scale feature fusion method is used to extract features from the input third feature map to obtain high-level semantic features and low-level spatial features; for the high-level semantic features and low-level spatial features, a feature fusion module is used for weighted fusion to generate a fourth feature map containing multi-scale information.

[0082] As a specific implementation, the multi-scale convolutional network is a deep learning architecture that extracts image features through convolutional kernels of different sizes. High-level semantic features contain object shape and category information, and low-level spatial features contain edge and texture details. In apple recognition, high-level features can identify the overall contour of the apple, and low-level features can capture the texture and color changes on the apple surface. For example, for a typical apple image, small-sized convolutional kernels can detect the fine texture of the peel, and large-sized convolutional kernels can identify the overall shape.

[0083] The feature fusion module combines different-level features in a weighted manner. By setting weight coefficients, the importance of high- and low-level features is adjusted. When the apple contour is clear, the weight of the high-level features can be set larger, such as 0.7; when relying on surface details, the weight of the low-level features can be increased to 0.5. Adaptive pooling unifies features of different scales to the same size for subsequent processing. For an input image of 256×256, the feature map can be normalized to 64×64.

[0084] Noise suppression uses filtering algorithms, commonly Gaussian filtering or median filtering. Gaussian filtering is suitable for dealing with uneven illumination, and median filtering is suitable for removing salt-and-pepper noise. In practical applications, if the apple image is strongly illuminated by sunlight to form a high-brightness area, a Gaussian filter with a σ value of 1.5 can be used for smoothing. For the case of being blocked by branches and leaves, a median filter with a radius of 3 can be used to remove interference.

[0085] Implementable, according to the fourth feature map, the neck module of the neural network is reconstructed, and a feature fusion mechanism is adopted to integrate feature information of different scales to obtain a fused fifth feature map, including:

[0086] The fourth feature map is optimized using a reconstruction method, and the structure of the neural network neck module is adjusted; for the multi-scale feature information in the fourth feature map, a feature fusion mechanism is designed to achieve the integration of features at different scales; through the information integration process, the fused fifth feature map is obtained.

[0087] As a specific implementation, the reconstruction method mainly includes two ways: feature pyramid reconstruction and cross-level feature reconstruction. Feature pyramid reconstruction adjusts feature maps at different scales to the same size through upsampling and downsampling operations to achieve multi-scale fusion of features. For example, in apple recognition, low-level edge texture features can be fused with high-level semantic features to effectively improve the accuracy of object detection.

[0088] The optimization of the neck module structure adopts a feature enhancement mechanism, and important features are highlighted through the combination of spatial attention and channel attention. Specifically, a dual-branch structure can be designed, where one branch is responsible for feature enhancement in the spatial dimension, and the other branch is responsible for feature enhancement in the channel dimension. Finally, the features of the two branches are adaptively weighted and fused.

[0089] The multi-scale feature fusion mechanism adopts a pyramid pooling structure to perform hierarchical processing on features at different scales. For example, the input feature map is respectively subjected to four different-scale pooling operations to obtain four feature maps with different receptive fields, and then they are aligned to the same scale through upsampling, and finally feature fusion is performed to generate a feature map containing multi-scale information.

[0090] Implementable, during the model training process, an improved loss function is adopted, combined with the problems of target boundary blur and scale change, to calculate the loss value between the prediction result and the true annotation, and an optimized loss value is obtained, including:

[0091] Analyze the features of target boundary blur and scale change, extract the blur degree of the target boundary and the key points of scale change; correct the deviation of the prediction result in combination with the target boundary information, and adjust the influence of scale change, and recalculate the loss value; adjust the weight of the loss function according to the correlation between target boundary blur and scale change to optimize the calculation process of the loss value; dynamically adjust the parameters of the loss function through fuzzy problem analysis and change problem feature extraction; until the maximum number of iterations is reached, an optimized loss value is obtained.

[0092] As a specific implementation, the improvement of the loss function first needs to consider the problem of blurred object boundaries. By analyzing the gradient of the boundary pixel distribution and combining the Gaussian blur coefficient, the boundary region is evaluated. For example, the boundary blur degree is divided into five levels. The weight of the region with the highest blur degree is set to 0.8, and the lowest is set to 0.2. When dealing with the problem of scale change, the relative size change of the object in the image needs to be considered. In response to this situation, an adaptive feature pyramid structure can be adopted to calculate the loss value for objects of different scales separately. For example, when the object size is less than 32 pixels, the overall contour features are mainly concerned, and the loss weight is 0.6; when the size is greater than 128 pixels, the weight of the detailed features is increased to 0.4. The blurred object boundaries and scale changes are often correlated. A joint weight matrix can be constructed to consider the influence of both factors simultaneously. By setting the joint weight coefficient, such as reducing the localization weight of blurred and small-scale objects to 0.3, while keeping the localization weight of clear and large-scale objects at 0.7. The improved loss function verifies the effect through comprehensive evaluation parameters. By adjusting the weight parameters, such as setting the boundary blur degree weight to 0.5 and the scale adaptation weight to 0.5, the detection accuracy of the model for objects in different states can be improved. Experiments show that the improved loss function can increase the detection accuracy by 15% and reduce the false detection rate by 20%.

[0093] Implementable, such as Figure 3 As shown, through the backpropagation algorithm, the optimized loss value is passed into the deep neural network to update the model parameters, obtaining a trained apple recognition model, including:

[0094] Through the backpropagation algorithm, the optimized loss value is passed into the deep neural network to update the model parameters, obtaining a trained apple recognition model. According to the training results of the apple recognition model, the blurred object boundary information is extracted to determine the blur degree, and the weight parameters of the loss function are adjusted. The blurred object boundary information is used to correct the deviation of the prediction result and recalculate the loss value to obtain the optimized loss value. The scale change features are extracted from the optimized loss value to determine the scale deviation and adjust the scale parameters of the model. According to the scale change features, the scale deviation of the prediction result is corrected and the loss value is recalculated to obtain the final loss value. The model parameters of the neural network are updated through the final loss value to complete the training of the apple recognition model. According to the trained apple recognition model, the prediction of the test data is performed to obtain the recognition result.

[0095] As a specific implementation, the backpropagation algorithm updates the weights and bias values layer by layer from back to front by calculating the gradients of the loss function with respect to the neural network parameters of each layer. Taking the apple recognition model as an example, when the model predicts an apple image, problems such as blurred boundaries and varying object scales may occur. For example, when recognizing apples in an orchard, due to factors such as lighting and occlusion, the boundaries of the apples will appear blurred. At this time, the weight adjustment of the loss function is required to improve the recognition effect. For the problem of blurred object boundaries, it can be quantitatively analyzed by setting a blur threshold. Suppose the blur degree is divided into five levels, with the first level to the fifth level indicating increasingly blurred boundaries. When it is detected that the blur degree of a certain area exceeds the third level, the weight of this area in the loss function needs to be increased. For example, in a certain picture, the boundary blur degree of an apple is level four, then the weight of its loss value is adjusted to twice the normal weight to strengthen the learning of the blurred area. When dealing with scale-varying features, the size changes of apples at different distances and angles need to be considered. For example, in the vision system of an orchard picking robot, apples nearby may occupy 30% of the image area, while apples in the distance may only occupy 5% of the area. At this time, multi-scale features are extracted through a feature pyramid network, and different weight coefficients are assigned to the prediction results of different scales. The influence of both boundary blur and scale change needs to be comprehensively considered during the calculation of the loss value.

[0096] Taking a certain training as an example, when an apple target with a boundary blur degree of level four and an area ratio of 20% is recognized, first, the loss weight is increased to twice according to the blur degree, then an appropriate feature layer is selected for prediction according to the area ratio, and finally, the optimized loss value is obtained. In the model parameter update process, the learning rate needs to be adjusted according to the gradient values of different layers. For example, in the convolutional layer dealing with the boundary blur problem, the learning rate can be appropriately increased to accelerate parameter convergence; while in the feature pyramid layer dealing with scale change, a smaller learning rate is required to maintain stability. This targeted parameter adjustment can improve the overall performance of the model. Finally, in the testing stage, the model can adapt to apple recognition tasks in different scenarios. For example, during orchard inspection, even when encountering blurred boundaries caused by cloudy days or apples at different distances, the optimized model can still maintain a high recognition accuracy. Through such multi-faceted optimization, the finally trained apple recognition model has strong practicality and robustness.

[0097] Implementably, the trained apple recognition model is deployed in the orchard environment to obtain orchard images in real time, input them into the model, judge the position and category of apples in the images, and output the recognition results, including:

[0098] Obtain real-time image data in the orchard environment. Use image processing technology to preprocess the images, removing noise and interference information to obtain clear images. Input the preprocessed images into the apple recognition model. Use the convolutional neural network algorithm to extract image features, determine the position and category of the apples, and generate a preliminary recognition result. According to the preliminary recognition result, use the object detection algorithm to correct the position. If the position deviation exceeds the preset threshold, adjust the bounding box to obtain accurate position information. Use the classification algorithm to conduct a secondary verification of the apple category. If the category confidence level is lower than the preset threshold, re-extract the features to determine the final category. According to the accurate position information and the final category, generate a complete recognition result. Use image annotation technology to mark the result on the image to obtain an annotated image. Transmit the annotated image to the storage system. Use data compression technology to reduce the storage space, ensuring data integrity and accessibility. According to the annotated image data in the storage system, use data analysis technology to count the number and distribution of apples, and generate an orchard apple distribution report.

[0099] As a specific implementation method, for obtaining real-time image data, it is necessary to reasonably configure the acquisition equipment. In the orchard environment, a high-definition camera array can be selected and arranged at the highest points in each area of the orchard. The camera focal length is set to three to five meters to cover the main growth areas of the fruit trees. The image preprocessing technology uses Gaussian filtering to remove random noise, median filtering to eliminate salt-and-pepper noise, and image enhancement to improve the contrast, ensuring the accuracy of subsequent recognition. The convolutional neural network feature extraction uses a multi-layer convolutional structure. The shallow network extracts basic features such as the edges and textures of the apples, and the deep network extracts semantic features such as the shape and color. The object detection algorithm uses the bounding box regression method, and sets the position deviation threshold to twenty pixels. When the detected apple position exceeds the threshold range, precise positioning is achieved by fine-tuning the bounding box coordinates.

[0100] The classification algorithm uses a multi-layer perceptron structure. The input feature vector dimension is 1280, the number of hidden layer nodes is 512, and the output is the probability distribution of the apple variety categories. The confidence threshold is set to 0.85. When it is lower than the threshold, re-extract the features and combine with context information for verification to improve the classification accuracy. For apples growing on the edge of the tree crown or blocked by branches and leaves, the multi-angle feature fusion method is used to improve the recognition effect. Image annotation uses a rectangular box to mark the apple position, and different colors are used to distinguish the varieties. The red box represents Fuji apples, and the yellow box represents Golden Delicious apples. The annotation information includes key information such as position coordinates, variety category, and confidence value. Data compression uses a lossless compression algorithm to control the compression ratio of the original image at about 5:1, which not only ensures the image quality but also saves storage space. The data analysis system classifies and stores the collected annotation data by time and region, establishes an orchard electronic map, and uses different colors to identify the apple distribution density. By comparing the distribution data of different periods, the growth status of the fruit trees can be analyzed and the yield can be predicted.

[0101] In dense areas, the system automatically marks the picking suggestion path to improve the picking efficiency. For fruit trees with abnormal growth and development, the system conducts comparative analysis based on historical data and promptly reminds fruit farmers to carry out pest control or nutritional regulation. In practical applications, this system can monitor the entire growth process of apples. Taking a demonstration park as an example, during the fruit growth period, the system collects canopy images every day. Through comparative analysis, it is found that the growth rate of apples is the fastest when the temperature is between 20 and 25 degrees Celsius, and the fruits on the south-facing canopy with sufficient sunlight grow 15% faster than those on the north-facing canopy. The system can also identify the coloring degree of the fruits to help fruit farmers grasp the best picking time. Through long-term data accumulation, a complete orchard management knowledge base is formed to provide data support for scientific planting.

[0102] Embodiment 2

[0103] As Figure 4 shown, this embodiment also provides an apple recognition system in a complex environment for implementing the method described above. The system includes:

[0104] An image preprocessing module for obtaining the original image containing apples in the orchard environment and preprocessing the original image using image enhancement technology to obtain the enhanced first image;

[0105] A labeling generation module for accurately labeling the position and category of each apple according to the apple area in the first image using a preset labeling tool to generate a second image containing target position and category information;

[0106] A feature extraction module for inputting the second image into a deep neural network, introducing an attention mechanism module in the feature extraction module, and obtaining a third feature map with attention weights by calculating the attention weights of the apple area;

[0107] A multi-scale feature fusion module for generating a fourth feature map containing multi-scale information based on the third feature map by using the multi-scale feature fusion method in the feature extraction module and combining high-level semantic information and low-level spatial details;

[0108] A feature reconstruction module for reconstructing the neck module of the deep neural network according to the fourth feature map, using a feature fusion mechanism to integrate feature information of different scales, and obtaining a fused fifth feature map;

[0109] A loss calculation module for training the deep neural network based on the fused fifth feature map. During the model training process, an improved loss function is used to calculate the loss value between the prediction result and the true label in combination with the problems of target boundary blur and scale change, and an optimized loss value is obtained;

[0110] A model training module, which is used to transmit the optimized loss value into the neural network through the backpropagation algorithm, update the model parameters, and obtain a trained apple recognition model;

[0111] An identification output module, which is used to deploy the trained apple recognition model into the orchard environment, obtain orchard images in real time, input them into the model, judge the position and category of apples in the images, and output the recognition results.

[0112] Embodiment III

[0113] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method.

[0114] Embodiment IV

[0115] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method are implemented.

[0116] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An apple recognition method under complex environments, characterized in that Including the following steps: Obtain the original image containing apples, and preprocess the original image using image enhancement technology to obtain the enhanced first image; According to the apple regions in the first image, use a preset annotation tool to annotate the position and category of each apple, generating a second image containing target position and category information; Input the second image into a deep neural network, introduce an attention mechanism module in the feature extraction module, calculate the attention weights of the apple regions, and obtain a third feature map with attention weights; Based on the third feature map, adopt a multi-scale feature fusion method in the feature extraction module, combine high-level semantic information and low-level spatial details, and generate a fourth feature map containing multi-scale information; According to the fourth feature map, reconstruct the neck module of the deep neural network, adopt a feature fusion mechanism to integrate feature information of different scales, and obtain a fused fifth feature map; Train the deep neural network based on the fused fifth feature map. During the model training process, adopt an improved loss function, combine the problems of target boundary blur and scale change, calculate the loss value between the prediction result and the true annotation, and obtain an optimized loss value; Through the backpropagation algorithm, transfer the optimized loss value into the deep neural network, update the model parameters, and obtain a trained apple recognition model; Deploy the trained apple recognition model to the orchard environment, obtain orchard images in real time, input them into the completed apple recognition model, judge the position and category of apples in the images, and output the recognition results.

2. The method according to claim 1, wherein The obtaining of the original image containing apples and preprocessing the original image using image enhancement technology to obtain the enhanced first image includes: Obtain the original image containing apples in the orchard environment. For the foliage occlusion and background interference in the original image, use the histogram equalization algorithm to enhance the image and obtain the enhanced first image.

3. The method according to claim 2, wherein The accurately annotating the position and category of each apple according to the apple regions in the first image using a preset annotation tool to generate a second image containing target position and category information includes: Adopt preset rules to divide the region of the first image and determine the distribution range of apple targets in the image; According to the target position information, call the annotation tool to draw bounding boxes for the apple regions; If the target category information matches the preset category, perform category marking on the pixel points within the annotation box; Optimize the boundaries of the annotation boxes through a preset threshold, obtain the optimized annotation box information, and then generate a second image containing target position and category information.

4. The method according to claim 1, wherein The inputting the second image into a deep neural network, introducing an attention mechanism module in the feature extraction module, and obtaining a third feature map with attention weights by calculating the attention weights of the apple regions includes: Use a deep neural network to extract features from the second image, embed an attention mechanism module in the feature extraction module; calculate the attention weight values of the apple regions, and simultaneously suppress the interference of the background domain and the foliage domain to obtain a third feature map with attention weights.

5. The method according to claim 4, characterized in that, Based on the third feature map, a multi-scale feature fusion method is adopted in the feature extraction module to combine high-level semantic information and low-level spatial details to generate a fourth feature map containing multi-scale information, including: In the feature extraction module, a multi-scale feature fusion method is used to extract features from the input third feature map to obtain high-level semantic features and low-level spatial features; For the high-level semantic features and low-level spatial features, a feature fusion module is used for weighted fusion to generate a fourth feature map containing multi-scale information.

6. The method according to claim 5, wherein According to the fourth feature map, the neck module of the neural network is reconstructed. A feature fusion mechanism is adopted to integrate feature information of different scales to obtain a fused fifth feature map, including: The fourth feature map is optimized by using a reconstruction method to adjust the structure of the neck module of the neural network; For the multi-scale feature information in the fourth feature map, a feature fusion mechanism is designed to realize the integration of features of different scales; through the information integration process, a fused fifth feature map is obtained.

7. The method according to claim 1, wherein During the model training process, an improved loss function is adopted to combine the problems of target boundary blur and scale change, calculate the loss value between the prediction result and the true annotation, and obtain an optimized loss value, including: Analyze the characteristics of target boundary blur and scale change, and extract the blur degree of the target boundary and the key points of scale change; Combine the target boundary information to correct the deviation of the prediction result and adjust the influence of scale change, and recalculate the loss value; According to the correlation between target boundary blur and scale change, adjust the weight of the loss function to optimize the calculation process of the loss value; Dynamically adjust the parameters of the loss function through fuzzy problem analysis and change problem feature extraction; Until the maximum number of iterations is reached, an optimized loss value is obtained.

8. An apple recognition system under complex environments, characterized in that For implementing the method according to any one of claims 1-7, the system includes: An image preprocessing module for obtaining an original image containing apples in an orchard environment, and preprocessing the original image by using an image enhancement technique to obtain a first enhanced image; A label generation module for accurately labeling the position and category of each apple according to the apple region in the first image by using a preset labeling tool to generate a second image containing target position and category information; A feature extraction module for inputting the second image into a deep neural network, introducing an attention mechanism module in the feature extraction module, and obtaining a third feature map with attention weights by calculating the attention weights of the apple region; A multi-scale feature fusion module for, based on the third feature map, adopting a multi-scale feature fusion method in the feature extraction module to combine high-level semantic information and low-level spatial details to generate a fourth feature map containing multi-scale information; A feature reconstruction module for, according to the fourth feature map, reconstructing the neck module of the deep neural network, adopting a feature fusion mechanism to integrate feature information of different scales to obtain a fused fifth feature map; A loss calculation module is used to train the deep neural network based on the fused fifth feature map. During the model training process, an improved loss function is adopted to calculate the loss value between the prediction result and the true annotation in combination with the problems of target boundary blurring and scale variation, and an optimized loss value is obtained; A model training module is used to transfer the optimized loss value into the neural network through the backpropagation algorithm, update the model parameters, and obtain a trained apple recognition model; An identification output module is used to deploy the trained apple recognition model to the orchard environment, obtain orchard images in real time, input them into the model, judge the position and category of apples in the images, and output the recognition results.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.