Fruit target detection method
By employing a dual-branch processing method and the Sobel operator to extract edge features, combined with a specific loss function to train the model, the problem of low fruit target detection accuracy of the YOLO model in complex agricultural environments was solved, achieving high-precision detection and improved generalization ability.
Patent Information
- Application Number
- CN202510831462.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, the YOLO model has low accuracy in detecting fruit targets when using limited computing resources in complex agricultural environments, and cannot adapt to the detection tasks of targets such as occluded fruits, small fruits, and densely growing fruits.
A dual-branch processing method is adopted, which compresses the spatial features of the target image through max pooling or average pooling, and extracts edge features using the Sobel operator to generate a comprehensive feature map. The model is trained by combining the loss functions of residual structure, attention component, distance penalty component, and moving average strategy component to improve detection accuracy.
Achieve high-precision detection of fruit targets in complex agricultural environments with low computing resources, enhance the model's generalization ability and adaptability to multiple environments, and reduce computational load while maintaining detection accuracy.
Smart Images

Figure CN120976514A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of agricultural environment target detection, in particular, the present application relates to a fruit target detection method. BACKGROUND
[0002] In the field of agriculture, it is necessary to perform target detection tasks autonomously in an agricultural environment by using edge devices such as automatic harvesting robots. For example, in the scenario of automatic picking of tomatoes, an automatic harvesting robot needs to obtain target images in the environment and perform tomato detection tasks to quickly and accurately identify occluded fruits, small fruits, and densely grown fruits in the target images under the interference of factors such as light changes, occlusions, and fruit density in the agricultural environment.
[0003] In the prior art, edge devices are used to carry YOLO and other models to perform target detection tasks in an agricultural environment. For example, by introducing a SimAM attention mechanism into the backbone network of a YOLOv8 model to extract feature information of a target image, using a detection structure Slim-Neck in the neck detection module of the YOLOv8 model, setting a loss function SIoU, and constructing a target detection model based on an improved YOLOv8, the detection accuracy and generalization recognition of the target in the target image are improved while ensuring the original detection rate of the YOLOv8.
[0004] However, since the network structure of the existing YOLO model requires a large amount of computing resources to ensure the detection accuracy and detection rate of target detection when performing target detection tasks in complex environments, the target detection accuracy in complex environments is greatly reduced in the case of limited computing resources, resulting in the fact that the YOLO model and other models in the prior art cannot adapt to performing target detection tasks in complex agricultural environments when deployed on edge devices such as automatic harvesting robots with limited computing power, and cannot ensure the detection accuracy of edge devices in performing target detection tasks on occluded fruits, small fruits, and densely grown fruits in complex agricultural environments.
[0005] As can be seen from the above, the prior art has the problem of low fruit target detection accuracy in complex agricultural environments under the premise of limited computing resources, which needs to be solved. SUMMARY
[0006] The present application provides a fruit target detection method in the technical field of agricultural environment target detection, which can solve the problem of low fruit target detection accuracy in complex agricultural environments under the premise of limited computing resources in the related art. The technical solution is as follows:
[0007] According to one aspect of the present application, a fruit target detection method comprises:
[0008] acquire a target image containing a fruit target, and perform first branch processing on the target image to compress spatial features of the target image, to generate a first pooling feature map, wherein the first branch processing is maximum pooling processing or average pooling processing;
[0009] perform second branch processing based on a Sobel operator to extract edge features of the target image, to generate a first edge feature map, wherein the second branch processing is edge detection processing;
[0010] perform image stitching on the first pooling feature map and the first edge feature map, to generate a comprehensive feature map;
[0011] input the comprehensive feature map into a pre-trained target recognition model to perform target recognition, to generate an identification result indicating a target prediction box of a fruit target in the comprehensive feature map.
[0012] According to an aspect of the present application, a target detection device comprises:
[0013] An image acquisition module is configured to acquire a target image containing a fruit target;
[0014] A first image processing module is configured to compress spatial features of the target image to generate a first pooling feature map, extract edge features of the target image based on a Sobel operator to generate a first edge feature map, and perform image stitching on the first pooling feature map and the first edge feature map to generate a comprehensive feature map;
[0015] A second image processing module is configured to perform convolution operation on the comprehensive feature map to generate a first convolution feature map, extract edge features in the comprehensive feature map based on a Sobel operator to generate a second edge feature map, and perform image fusion on the comprehensive feature map, the first convolution feature map, and the second edge feature map based on a residual structure, to update the comprehensive feature map with a second fused image;
[0016] A third image processing module is configured to perform average pooling processing on the comprehensive feature map, and perform image segmentation to obtain a first feature subgraph and a second feature subgraph, perform maximum pooling processing on the first feature subgraph to generate a second pooling feature map, perform convolution operation on the second feature subgraph to generate a third convolution feature map, and perform image stitching on the second pooling feature map and the third convolution feature map to update the comprehensive feature map with a fourth fused image;
[0017] The target recognition module is configured to acquire a target recognition model trained in advance based on a loss function comprising an attention component, a distance penalty component and a moving average strategy component, and input the integrated feature map into the pre-trained target recognition model to perform target recognition and generate an identification result of a target prediction box indicating a fruit target in the integrated feature map.
[0018] In an example embodiment, the Sobel operator comprises a horizontal edge convolution kernel and a vertical edge convolution kernel; and the first image processing module comprises:
[0019] a horizontal convolution subunit configured to perform a convolution operation on the target image and the horizontal edge convolution kernel to extract edge features in a horizontal direction of the target image and generate a horizontal edge feature map;
[0020] a vertical convolution subunit configured to perform a convolution operation on the target image and the vertical edge convolution kernel to extract edge features in a vertical direction of the target image and generate a vertical edge feature map;
[0021] an image fusion subunit configured to perform image fusion on the horizontal edge feature map and the vertical edge feature map to generate a first edge feature map.
[0022] In an example embodiment, the first image processing module further comprises:
[0023] a first channel adjustment subunit configured to perform channel adjustment on the first pooled feature map and the first edge feature map to adjust the first pooled feature map and the first edge feature map to a first set number of channels;
[0024] a concatenation subunit configured to perform channel concatenation on the first pooled feature map and the first edge feature map along a channel dimension based on the first set number of channels to generate a first fused image;
[0025] a second channel adjustment subunit configured to perform a convolution operation on the first fused image to adjust a number of channels of the first fused image to a second set number of channels to generate an integrated feature map, wherein the second set number of channels is less than the first set number of channels.
[0026] In an example embodiment, the target detection device further comprises:
[0027] a first convolution module configured to perform a convolution operation on the integrated feature map to generate a first convolution feature map;
[0028] an edge feature extraction module configured to extract edge features in the integrated feature map based on a Sobel operator to generate a second edge feature map;
[0029] The first image fusion module is configured to perform image fusion on the integrated feature map, the first convolution feature map and the second edge feature map based on a residual structure to update the integrated feature map with a generated second fusion image.
[0030] In an example embodiment, the first image fusion module comprises:
[0031] A first image fusion unit is configured to perform image fusion on the first convolution feature map and the second edge feature map to generate a third fusion image.
[0032] A second convolution unit is configured to perform convolution operation on the third fusion image to generate a second convolution feature map.
[0033] A second image fusion unit is configured to connect the integrated feature map and the second convolution feature map based on a residual structure and perform image fusion to generate a second fusion image.
[0034] An updating unit is configured to update the integrated feature map with the generated second fusion image.
[0035] In an example embodiment, the target detection device further comprises:
[0036] An average pooling module is configured to perform average pooling processing on the integrated feature map to generate an average pooling feature map.
[0037] An image segmentation module is configured to segment the average pooling feature map into a first feature submap and a second feature submap which is different from the first feature submap in image channel dimension.
[0038] A maximum pooling module is configured to perform maximum pooling processing on the first feature submap to generate a second pooling feature map.
[0039] A third convolution module is configured to perform convolution operation on the second feature submap to generate a third convolution feature map.
[0040] A second image fusion module is configured to perform image splicing on the second pooling feature map and the third convolution feature map to update the integrated feature map with a generated fourth fusion image.
[0041] In an example embodiment, the target recognition model is trained based on a loss function comprising an attention component, a distance penalty component and a moving average strategy component; and the target detection device further comprises:
[0042] A training module is configured to obtain a pre-trained model and select a training set to perform target recognition training on the target detection model, wherein the training pictures in the training set comprise fruit targets and corresponding real boundary boxes.
[0043] an attention module configured to generate an attention result based on an attention component calculating an overlap between the real bounding box and a target prediction box generated by the target detection model during the training process;
[0044] a distance penalty module configured to generate a distance penalty result based on a distance penalty component calculating a relative distance difference between the real bounding box and the target prediction box during the training process;
[0045] a parameter updating module configured to calculate a loss intensity based on a moving average strategy component for the attention result and the distance penalty result, and to update model parameters of the target detection model based on the loss intensity until the target recognition training is completed.
[0046] In an example embodiment, the attention module comprises:
[0047] an overlap parameter unit configured to calculate an overlap between the real bounding box and a target prediction box generated by the target detection model, and to generate an overlap degree parameter;
[0048] a weight unit configured to obtain an attention parameter and to set a weight of the target prediction box based on a difference between the overlap degree parameter and the attention parameter;
[0049] an attention result unit configured to generate an attention result based on the weight of the target prediction box.
[0050] In an example embodiment, the parameter updating module comprises:
[0051] a training stage acquisition unit configured to acquire a training stage corresponding to each loss intensity, the training stage being pre-divided based on a model performance training requirement of the target recognition training;
[0052] a training stage determination unit configured to determine a training stage of the target recognition training based on a model performance change trend of the target recognition model calculated based on the attention result, the distance penalty result and the moving average strategy component;
[0053] a parameter updating unit configured to update model parameters of the target recognition model based on a loss intensity corresponding to the training stage until the target recognition training is completed.
[0054] In an example embodiment, the target recognition module comprises:
[0055] an image processing network unit configured to input the comprehensive feature map into a hierarchically iteratively connected image processing network for image processing, wherein each level of the image processing network performs first branch processing and second branch processing on a target image based on a comprehensive feature map output by a previous level, and outputs a comprehensive feature map corresponding to the level.
[0056] a first image scale unit configured to determine an image scale of a comprehensive feature map output by each level of the image processing network in the image processing process;
[0057] a second image scale unit configured to determine an image scale of target recognition based on a size of the fruit target in the target image;
[0058] a recognition result generation unit configured to input the comprehensive feature map corresponding to the image scale into the target recognition model to perform target recognition and generate a recognition result.
[0059] The technical scheme provided in the present application has the following beneficial effects:
[0060] In the above technical scheme, the target image is subjected to pooling processing in a double-branch form and edge detection based on a Sobel operator to compress the spatial features of the target image, extract the edge features of the target image, generate a comprehensive feature map, input the comprehensive feature map into a target recognition model to perform target recognition, enhance the ability to capture edge information of the target image, improve the detection accuracy of small-size and partially occluded fruit targets, reduce the resolution of the comprehensive feature map, and thus reduce the target recognition calculation amount, while retaining the overall structure and background information of the image to ensure the recognition accuracy. Thus, high-precision detection of fruit targets in a complex agricultural environment can be realized while maintaining low calculation cost, and the model can be deployed on edge devices, and the generalization ability and multi-environment adaptability of the model are improved. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical schemes in the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced.
[0062] Figure 1 is a schematic diagram according to the implementation environment involved in the present application;
[0063] Figure 2 is a flowchart of a fruit target detection method according to an exemplary embodiment;
[0064] Figure 3 is a structural schematic diagram of a fruit target detection structure according to an exemplary embodiment;
[0065] Figure 4 is a structural schematic diagram of a fruit target detection structure according to an exemplary embodiment;
[0066] Figure 5 is a structural schematic diagram of a fruit target detection structure according to an exemplary embodiment;
[0067] Figure 6is a structural schematic diagram of a target prediction frame according to an example embodiment;
[0068] Figure 7 is a specific implementation schematic diagram of a fruit target detection method in an application scenario;
[0069] Figure 8 is a device schematic diagram of a fruit target detection method in an application scenario. DETAILED DESCRIPTION
[0070] The embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar notations represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as a limitation of the present application.
[0071] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an" and "the" used herein also include the plural forms. It should be further understood that the use of the phrase "comprising" in the specification of the present application means that a feature, integer, step, operation, element and / or component exists, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.
[0072] As before, there is still a problem of low fruit target detection accuracy in the prior art under the premise of using limited computing resources in a complex agricultural environment.
[0073] To this end, the fruit target detection method provided by the present application can effectively realize high-precision detection of fruit targets in a complex agricultural environment while maintaining low computing cost, and has the ability to be deployed on edge devices, while improving the generalization ability and multi-environment adaptability of the model. Accordingly, the above method is suitable for a fruit target detection device, which can be deployed on an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a notebook computer, a server, etc. The electronic device can be an electronic device with a central control function, for example, the electronic device includes a gateway, etc. The electronic device can also be a portable mobile electronic device, for example, the electronic device includes a smart phone, a tablet computer, etc.
[0074] In order to make the purposes, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0075] Figure 1 A schematic diagram of an implementation environment related to a fruit target detection method. It should be noted that this implementation environment is only an example suitable for the present application and should not be considered as providing any limitation on the scope of use of the present application.
[0076] The implementation environment includes a collection end 110 and a service end 130.
[0077] Specifically, the collection end 110, which can also be regarded as a target image collection device, includes but is not limited to a camera, a smartphone, and other electronic devices with target image collection functions. For example, the collection end 110 is a mobile agricultural robot equipped with a camera.
[0078] The service end 130 can be a desktop computer, a notebook computer, a server, or other electronic devices. It can also be a computer cluster composed of multiple servers, or even a cloud computing center composed of multiple servers. The service end 130 is used to provide background services, such as fruit target detection services.
[0079] The service end 130 and the collection end 110 are pre-established in network communication connection through wired or wireless means, and the data transmission between the service end 130 and the collection end 110 is realized through the network communication connection. The transmitted data includes but is not limited to target images and the like.
[0080] In an application scenario, through the interaction between the collection end 110 and the service end 130, the collection end 110 samples the agricultural environment and obtains a target image uploaded to the service end 130 to request the service end 130 to provide fruit target detection services.
[0081] For the service end 130, after receiving the target image uploaded by the collection end 110, the fruit target detection service is called to generate a comprehensive feature map through low-computing-cost image processing based on the obtained target image, and to identify the fruit target in the comprehensive feature map through a target recognition model to generate an identification result of a target prediction box indicating the fruit target in the comprehensive feature map. In this way, the problem of low fruit target detection accuracy in a complex agricultural environment under the premise of limited computing resources in the related art is solved.
[0082] Please refer to Figure 2 The embodiments of the present application provide a fruit target detection method, which is suitable for electronic devices such as desktop computers, notebook computers, servers, and the like.
[0083] In the following method embodiment, for the convenience of description, the execution subject of each step of the method is taken as an electronic device for example, but it does not constitute a specific limitation.
[0084] As shown in the method can include the following steps: Figure 2
[0085] Step 210, obtaining a target image containing a fruit target, and performing first branch processing on the target image to compress the spatial features of the target image, and generating a first pooling feature map.
[0086] Among them, the first branch processing is maximum pooling processing or average pooling processing, and the target image is an original image obtained by shooting in an agricultural scene, for example, an image obtained by shooting the fruit target in the surrounding environment during the automatic driving process of an agricultural robot.
[0087] Specifically, after obtaining the target image, the corresponding pooling processing mode is selected based on the spatial features of the target image to perform first branch processing on the target image, so as to compress the spatial features of the target image, for example, according to the demand of target detection of the target image, it is determined whether it is necessary to retain the significant features in the target image or to smooth the target image as much as possible, thereby determining whether to use the maximum pooling processing to retain the significant features in the target image or the average pooling processing to smooth the features in the target image in the first branch processing, and the width and height of the target image and other spatial features are compressed through the first branch processing. Finally, a first pooling feature map with lower resolution than the input target image is generated, which retains the overall structure information in the target image.
[0088] Step 230, performing second branch processing based on the Sobel operator to extract the edge features of the target image to generate a first edge feature map.
[0089] Among them, the second branch processing is edge detection processing.
[0090] Specifically, the second branch processing is an image processing path independent of the first branch processing, and the first branch processing and the second branch processing of the target image can be performed at the same time in the image processing flow. While the first branch processing performs maximum pooling processing or average pooling processing on the target image, the second branch processing independently performs edge detection on the target image, that is, extracts edge features through the Sobel operator to highlight the edges in the target image, and finally constructs a first edge feature map through the extracted edge features.
[0091] In one possible implementation, the Sobel operator is composed of a horizontal convolution kernel and a vertical convolution kernel.
[0092] Step 250, image splicing the first pooling feature map and the first edge feature map to generate a comprehensive feature map.
[0093] Specifically, after the first branch processing and the second branch processing are both completed, the first pooled feature map and the first edge feature map generated for the target image in different dimensions can be obtained, and by image splicing the two, a comprehensive feature map combining the edge information in the first pooled feature map and the overall structural information in the first edge feature map is generated. By compressing the spatial features and retaining the overall structural information, and combining with the edge sensitive features in the edge feature map, the network can better isolate small objects and process areas with uneven illumination.
[0094] In a possible implementation, the image splicing includes splicing based on the feature map amplitudes and splicing based on the feature map channels.
[0095] In a possible implementation, the first branch processing, the second branch processing and the image splicing on the target image are performed through an SPStem structure, as shown in the following formula: Figure 3 After obtaining the target image of 640*640, the SPStem structure first performs convolution operation through the Conv convolution block to reduce the target image to 320*320, and then inputs the MaxPool2d module in a double-branch form to perform maximum pooling processing, and performs edge detection through the Sobel-Conv module to obtain the first pooled feature map and the first edge feature map. The two feature maps are fused through the Concat image connection block to generate a feature map containing edge information and overall structural information. The fused feature map may need to pass through two convolution layers to further integrate the feature information and generate the final comprehensive feature map.
[0096] In step 270, the comprehensive feature map is input into a pre-trained target recognition model to perform target recognition, and an identification result indicating a target prediction box of the fruit target in the comprehensive feature map is generated.
[0097] Specifically, after obtaining the comprehensive feature map, the target recognition model predicts the bounding box and the class probability of each region in the comprehensive feature map, and generates a corresponding target prediction box in the region where the fruit target exists.
[0098] In a possible implementation, the target recognition model is constructed based on the YOLO11 model.
[0099] In the above process, the resolution of the target image is reduced through the first branch processing for pooling, so that the calculation amount of the subsequent target recognition process is reduced, while the overall structural information and background information of the image can be retained to ensure the detection accuracy, and the edge features of the image are extracted through the second branch processing independent of the first branch processing. The edge information reflecting the characteristics of small targets and occluded targets is obtained to ensure the recognition accuracy of small targets and occluded targets in the subsequent target recognition process. Through the double-branch design, the target image is processed specifically to improve the processing efficiency and reduce the calculation resources required for image processing, so that the fruit target detection accuracy can be maintained under the premise of using limited calculation resources.
[0100] In an example embodiment, step 230 can include the following steps:
[0101] Step 231, performing convolution operation on the target image and a horizontal edge kernel to extract the edge features of the target image in the horizontal direction and generate a horizontal edge feature map.
[0102] The horizontal convolution kernel is an edge feature extraction structure specialized for the horizontal direction, and the convolution operation of the target image and the horizontal edge kernel enhances the horizontal edge response of the target image, quickly extracts the edge features in the specific horizontal direction while consuming less calculation amount, and generates the horizontal edge feature map.
[0103] Step 233, performing convolution operation on the target image and a vertical edge kernel to extract the edge features of the target image in the vertical direction and generate a vertical edge feature map.
[0104] The vertical convolution kernel is an edge feature extraction structure specialized for the vertical direction, and the convolution operation of the target image and the vertical edge kernel enhances the vertical edge response of the target image and generates the vertical edge feature map.
[0105] Step 235, performing image fusion on the horizontal edge feature map and the vertical edge feature map to generate a first edge feature map.
[0106] The image fusion of the horizontal edge feature map and the vertical edge feature map is to combine the horizontal direction edge features in the horizontal edge feature map with the vertical direction edge features in the vertical edge feature map to generate an edge feature map indicating the overall edge features.
[0107] In a possible implementation, the image fusion is based on the feature map amplitude or the feature map channel.
[0108] Through the above process, the edge features of the target image in a specific direction are extracted by the horizontal edge convolution kernel and the vertical edge convolution kernel, which can reduce the calculation amount and improve the calculation speed of edge feature extraction. The image fusion ensures that the overall edge features are included in the comprehensive feature map, ensuring the accuracy of the fruit target detection.
[0109] In an example embodiment, step 250 can include the following steps:
[0110] Step 251, channel adjustment is performed on the first pooled feature map and the first edge feature map to adjust the first pooled feature map and the first edge feature map to a first set number of channels.
[0111] In a possible implementation, the first set number of channels is set to the number of channels of the first edge feature map, and the number of channels of the first pooled feature map is adjusted to keep the number of channels of the first pooled feature map and the first edge feature map consistent, facilitating subsequent image fusion processing.
[0112] Step 253, based on the first set number of channels, the first pooled feature map and the first edge feature map are spliced along the channel dimension to generate a first fused image.
[0113] Specifically, by keeping the number of channels of the first pooled feature map and the first edge feature map the same, under the condition of a specified first set number of channels, the features extracted by different paths are integrated along the channel dimension, and the background and overall structure information in the first pooled feature map and the edge features in the first edge feature map are integrated to generate the first fused image.
[0114] Step 255, the first fused image is subjected to a convolution operation to adjust the number of channels of the first fused image to a second set number of channels to generate a comprehensive feature map.
[0115] Wherein, the second set number of channels is less than the first set number of channels.
[0116] Specifically, by additionally performing a convolution operation on the first fused image, the spliced feature map is further fused. This step aims to fuse the spliced multi-channel feature map to reduce the number of image channels, so that the calculation amount of subsequent image processing can be further reduced.
[0117] Through the above process, the first pooled feature map and the first edge feature map generate a comprehensive feature map with more rich semantic information and more compact representation, effectively integrating multi-path information, providing better feature input for subsequent processing modules, and reducing the consumption of calculation amount.
[0118] In an example embodiment, the fruit target detection method further includes the following steps:
[0119] Step 310, a convolution operation is performed on the integrated feature map to generate a first convolution feature map.
[0120] The first convolution feature map is obtained by performing a convolution operation on the integrated feature map through a standard convolution layer.
[0121] Step 330, edge features in the integrated feature map are extracted based on a Sobel operator to generate a second edge feature map.
[0122] The Sobel operator performs a second edge feature extraction on the basis of the integrated feature map generated by the first edge feature extraction, strengthens the edge information in the target image, and generates a second edge feature map with more prominent edge features.
[0123] Step 350, the integrated feature map, the first convolution feature map and the second edge feature map are image fused based on a residual structure to update the integrated feature map with a generated second fusion image.
[0124] In an exemplary embodiment, step 350 can include the following steps:
[0125] Step 351, the first convolution feature map and the second edge feature map are image fused to generate a third fusion image.
[0126] Specifically, the integrated feature map is processed by two different feature extraction methods of convolution operation and edge feature extraction, and the first convolution feature map and the second edge feature map with different feature dimensions are obtained. The first convolution feature map and the second edge feature map are image fused, the deep semantic features of the convolution feature map and the image edge information of the edge feature map are retained in the third fusion image, and the image representation ability is improved.
[0127] Step 353, a convolution operation is performed on the third fusion image to generate a second convolution feature map.
[0128] By additionally performing a convolution operation on the first fusion image, the fused and spliced feature map is further fused to generate a second convolution feature map with more rich semantic information and more compact representation, and effective integration of the multi-path information of the convolution operation and the edge feature extraction is realized.
[0129] Step 355, the integrated feature map and the second convolution feature map are connected based on a residual structure, and image fusion is performed to generate a second fusion image.
[0130] In a possible implementation, the second fusion image is generated by the SEDFF structure, and the SEDFF structure is as follows: Figure 4As shown, the 160*160 comprehensive feature map is respectively input into the Conv standard convolution block for convolution operation and the Sobel-Conv module for edge feature extraction, and the first image fusion is performed through the Concat image connection block. After the fused image is processed through the Conv convolution block, the original input comprehensive feature map is added to the processed fused feature map through the residual connection, to form a composite second fused image.
[0131] In step 357, the comprehensive feature map is updated with the generated second fused image.
[0132] In a possible implementation, after the SPStem structure is acquired, the SEDFF structure is constructed, and the comprehensive feature map generated by the SPStem structure is processed through the SEDFF structure to generate a new comprehensive feature map.
[0133] Through the above process, through further operations including convolution operation and edge feature extraction on the comprehensive feature map, the model's extraction and integration capability for edge information is further enhanced on the basis of the first edge feature extraction. Through the fusion of the edge feature map extracted by the Sobel operator and the feature map obtained by the traditional convolution operation, a more accurate edge feature map is generated. At the same time, through the residual connection, the key features of the comprehensive feature map image are retained, so that the performance of the second fused image will not lose the key features in the processing process, the edge information in the second fused image is enhanced, and the subsequent target recognition model's recognition accuracy for small targets and occluded targets is improved.
[0134] In an exemplary embodiment, the fruit target detection method further includes the following steps:
[0135] In step 410, the comprehensive feature map is subjected to average pooling processing to generate an average pooling feature map.
[0136] The spatial dimension of the comprehensive feature map is further reduced through the average pooling processing, the comprehensive feature map is down-sampled, and the features in the comprehensive feature map are smoothed.
[0137] In step 430, the average pooling feature map is divided into a first feature sub-map and a second feature sub-map which is different from the first feature sub-map in image channel dimension.
[0138] The average pooling feature map is divided into different sub-maps along the channel dimension, the first feature sub-map and the second feature sub-map are different in the number of channels, and the first feature sub-map and the second feature sub-map are respectively used for subsequent image processing of different granularities based on different numbers of channels.
[0139] In step 450, the first feature sub-map is subjected to maximum pooling processing to generate a second pooling feature map.
[0140] Step 470, a convolution operation is performed on the second feature subgraph to generate a third convolution feature map.
[0141] Specifically, the local features of the second feature subgraph are extracted through the convolution operation to obtain fine-grained local detail information and further reduce the spatial dimension of the feature map. The coarse-grained overall structure information is extracted through the max-pooling processing, and the feature information of different granularities in the first feature subgraph and the second feature subgraph is fully utilized.
[0142] Step 490, the second pooled feature map and the third convolution feature map are image spliced to generate a fourth fusion image to update the comprehensive feature map.
[0143] In one possible implementation, the second fusion image is generated by an ADown structure, as shown in Figure 5 As shown, after the comprehensive feature map is subjected to the average pooling processing by the AvgPool2d block, the image is segmented by the Spilt block, and different subgraphs are respectively input into the Conv convolution block for convolution operation and the MaxPool2d block for max-pooling processing. Meanwhile, the second pooled feature map generated after the max-pooling processing is subjected to channel number adjustment by the Conv convolution block, so that the channel numbers of the second pooled feature map and the third convolution feature map are consistent, so as to be fused along the channel dimension by the subsequent Concat block,
[0144] In one possible implementation, after the SEDFF structure is obtained, the ADown structure is constructed, and the comprehensive feature map generated by the ADown structure is subjected to down-sampling processing by the SEDFF structure to generate a new comprehensive feature map.
[0145] Through the above process, the spatial dimension of the feature map is reduced by the average pooling processing, and the feature map is further subjected to down-sampling and feature extraction by the convolution layer and the max-pooling, so as to fully utilize the background and overall structure information in the average pooling path and the local detail information in the convolution path. The information diversity in the down-sampling process of the comprehensive feature map is maintained, and the features obtained by different paths are integrated by the multi-path information integration to create more rich feature representations, which is helpful for capturing and processing multi-scale information, improves the feature capturing capability of the model, and thus realizes the down-sampling of the comprehensive feature map while reducing the complexity of the model, and reduces the calculation amount in the subsequent target recognition process.
[0146] In an exemplary embodiment, the training process of the target recognition model can include the following steps:
[0147] Step 510, a pre-trained model is obtained, and a training set is selected to train the target detection model for target recognition.
[0148] The training pictures in the training set include fruit targets and corresponding real boundary boxes. The target recognition model is trained based on the training set and a loss function including an attention component, a distance penalty component and a moving average strategy component.
[0149] At step 530, the attention component is used to calculate the overlap between the real boundary box and the target prediction box generated by the target detection model during the training process, and an attention result is generated.
[0150] In an exemplary embodiment, step 530 can include the following steps:
[0151] At step 531, the overlap between the real boundary box and the target prediction box generated by the target detection model is calculated, and an overlap parameter is generated.
[0152] In a more likely implementation, the overlap parameter is calculated as follows:
[0153]
[0154] where B pred is the target prediction box, B gt is the real boundary box, and IoU is the overlap parameter.
[0155] At step 533, the attention parameter is obtained, and based on the difference between the overlap parameter and the attention parameter, the weight of the target prediction box is set.
[0156] Based on the calculation of the overlap, by combining the overlap parameter with the principle of FocalLoss, replacing the original IoU with the IoU converted by Focaler, and introducing the attention parameter d and the attention parameter u set for the key samples, the weight of the target prediction box can be set. The attention component weight formula is as follows:
[0157]
[0158] At step 535, the attention result is generated based on the weight of the target prediction box.
[0159] The attention result indicates the weight of the target prediction box for a specific key sample. The closer the target prediction box is to the real boundary box related to the attention parameter d and the attention parameter, the higher the overlap is, and the greater the weight is.
[0160] Through the above process, by setting the weight and introducing the learnable attention parameter, the influence on the key samples is enhanced and the high-precision prediction is distinguished, and the recognition accuracy of the target recognition model is improved.
[0161] At step 550, the distance penalty component calculates a relative distance difference between the real bounding box and the target prediction box during the training process to generate a distance penalty result.
[0162] The distance penalty result is obtained by generating a normalized distance penalty term according to the relative distance difference.
[0163] Specifically, the target recognition model determines the boundary of the real bounding box and the target prediction box every time a target prediction box is generated, and calculates the relative distance between the boundary lines of the real bounding box and the target prediction box. The smaller the relative distance, the closer the target prediction box to the real prediction box, and the stronger the distance penalty result indicates the identification ability of the target recognition model for the fruit target.
[0164] For example, referring back to Figure 6 , the real bounding box target bbox in the figure has a length of Wgt and a width of hgt. The distance dw1, dw2, dh1, dh2 between the target prediction box anchor bbox and the real bounding box target bbox is calculated, and the coordinates of the real bounding box and the target prediction box are set as (b1x1, b1y1, b1x2, b1y2) and (b2x1, b2y1, b2x2, b2y2), respectively. At this time, the relative distance difference formula is as follows:
[0165] dw1 = |min(b1 x2 ,b1 x1 )-min(b2 x2 ,b2 x1 )|
[0166] dw2 = |max(b1 x2 ,b1 x1 )-max(b2 x2 ,b2 x1 )|
[0167] dh1 = |min(b1 y2 ,b1 y1 )-min(b2 y2 ,b2 y1 )|
[0168] dh2 = |max(b1 y2 ,b1 y1 )-max(b2 y2 ,b2 y1 )|
[0169] At this time, the normalized distance penalty term formula of the distance penalty component is as follows:
[0170]
[0171] At step 570, the loss intensity is calculated based on the attention result and the distance penalty result by the moving average strategy component, and the model parameter update is performed on the target detection model based on the loss intensity until the target recognition training is completed.
[0172] In a possible implementation, the loss intensity is determined by a loss term formula of PIoU2, and the calculation formula of the loss term of PIoU2 is as follows:
[0173] q=e -P
[0174]
[0175] wherein λ=1.3 is a scaling coefficient, providing more comprehensive spatial supervision than the center point-based index.
[0176] In a possible implementation, an adaptive scaling mechanism is introduced to construct the moving average strategy component:
[0177] The adaptive scaling mechanism formula of the moving average strategy component is as follows:
[0178]
[0179] The average IoU is updated by an exponential moving average, and the update formula is as follows:
[0180]
[0181] wherein m=0.01 and N is the number of positive samples in each small batch.
[0182] The loss function formula for training the target recognition model is as follows:
[0183]
[0184] wherein α=1.7 and δ=2.7 follow the design of the moving average strategy component.
[0185] Through the above process, the influence on key samples is enhanced and high-precision predictions are distinguished. The relative positions between the boundaries of the bounding boxes are considered by the distance penalty term, providing more comprehensive spatial guidance than the method based on the center point only, and providing comprehensive spatial supervision. Both the geometric penalty mechanism of the distance penalty component and the emphasis on high-quality predictions of the attention component are retained, the moving average strategy is introduced, the adaptive scaling is applied to enhance the PIoU2 loss, and the loss intensity is dynamically adjusted, so that the model can adapt to different training stages. This design not only improves the accuracy of target positioning, but also enhances the adaptability of the model to complex agricultural environments. Breakthroughs have been made in balancing spatial supervision, quality-aware training, and adaptive optimization, overcoming the limitations of traditional methods, and achieving target recognition with higher robustness.
[0186] In an example embodiment, step 570 can include the following steps:
[0187] Step 571, obtain the training phase corresponding to each loss intensity, and the training phase is divided in advance based on the model performance training requirement of the target recognition training.
[0188] In a possible implementation, the training phase includes a rapid decline phase, a stable phase and an overfitting phase. Wherein, the model performance training requirement of the basic pattern and feature in the rapid learning data corresponding to the training start phase divides the rapid decline phase, the model performance training requirement of the stable update model parameter corresponding to the training stable phase divides the stable phase, and the model performance training requirement of the training loss reduction corresponding to the training later phase divides the overfitting phase.
[0189] Step 573, calculate the model performance change trend of the target recognition model based on the attention result, the distance penalty result and the moving average strategy component, to determine the training phase of the target recognition training.
[0190] Specifically, based on the attention result, the distance penalty result and the moving average strategy component, the loss function change of the current target recognition model for the real prediction frame in the training set is predicted. In the beginning stage of training, the loss function rapidly decreases. In the stable phase, when the decline speed of the loss function slows down significantly, it enters a relatively stable state. In the overfitting phase, the training loss continues to decrease, but the verification loss begins to rise.
[0191] Step 575, update the model parameters of the target recognition model based on the loss intensity corresponding to the training phase until the target recognition training is completed.
[0192] In a possible implementation, the training corresponding to multiple epoch periods is performed in the rapid decline phase until the decline speed of the loss function begins to slow down.
[0193] In a possible implementation, the model parameters are updated in a small amplitude in the stable phase.
[0194] In a possible implementation, regularization means such as dynamic weight adjustment in WFPIoU loss function are used for training in the overfitting phase to alleviate the overfitting phenomenon in the training process.
[0195] Through the above process, by dynamically adjusting the loss intensity, the target recognition model can adapt to different stages of training. Not only the accuracy of target positioning is improved, but also the adaptability of the model to complex agricultural environment is enhanced.
[0196] In an example embodiment, 270 can include the following steps:
[0197] Step 271, input the integrated feature map into the image processing network with hierarchical iterative connection for image processing.
[0198] Each level in the image processing network processes the integrated feature map output by the previous level as the target image for first branch processing and second branch processing, and outputs the integrated feature map of the corresponding level.
[0199] Step 273, determine the image size of the integrated feature map output by each level of the image processing network during image processing.
[0200] Step 275, determine the image size of target recognition based on the size of the fruit target in the target image.
[0201] Step 277, input the integrated feature map corresponding to the image size into the target recognition model for target recognition to generate a recognition result.
[0202] Specifically, the image processing layer is constructed through the above-mentioned integrated feature map generation process, the integrated feature map output by each layer is taken as the target image of the next layer, and iterative connection of different levels is constructed. When the target image is input into the image processing network, each image processing layer will output the integrated feature map of the corresponding level. The higher the level, the lower the resolution of the integrated feature map, the larger the scale, and the more abstract the extracted features, which can capture more semantic information such as shape and category in the image. By determining the image size that best fits the size of the fruit target in the target image, the integrated feature map corresponding to the size is input into the target recognition model for target recognition, which can improve the accuracy of target recognition.
[0203] Through the above process, it is ensured that the image size of the integrated feature map conforms to the size of the fruit target, improving the accuracy of fruit target detection and improving the recognition efficiency.
[0204] Figure 7 is a specific implementation diagram of a fruit target detection method in an application scenario. In this application scenario, the SE-YOLO detection model is used for image processing and target detection of the target image. The SPStem module fuses the Sobel operator, which is used to directly extract the edge features of the image and perform maximum pooling processing or average pooling processing.
[0205] The SEDFF module is integrated into the Backbone of the SE-YOLO detection model, and the SEDFF module is used to replace the traditional C3k2 module in the backbone network of the SE-YOLO detection model. The SEDFF module further performs convolution operation and edge feature extraction on the comprehensive feature map output by the SPStem module, further enhancing the extraction and fusion of edge information. The SEDFF module emphasizes horizontal and vertical edge information and fuses the results to generate more accurate edge feature maps, thereby improving the model's ability to identify tomato boundaries.
[0206] The ADown module reduces the size of the feature map through average pooling and splits the feature map along the channel dimension for downsampling and feature extraction, respectively. Finally, the ADown module performs feature fusion to achieve efficient downsampling.
[0207] The 11Detect module is constructed based on the WFPIoU loss function, which calculates the difference between the predicted results and the true labels. The 11Detect module is used to predict the position and category of the target based on the feature map output by the SEDFF module during target identification. By dynamically adjusting the loss weight, the model's adaptability to different quality predictions is improved, thereby improving the positioning accuracy. The WFPIoU loss function combines the Focaler-IoU attention component, the Powerful-IoUv2 distance penalty component, and the Wise-IoUv3 moving average strategy component, which respectively emphasize high-quality samples, provide more comprehensive spatial supervision, and dynamically adjust the loss intensity. In addition, the 11Detect module also includes the BCE loss function, which is used for category prediction and measures the difference between the model's predicted probability distribution and the true label. The BCE loss function is used to optimize the prediction of the target category for each target prediction box in the target detection task, enabling the model to more accurately distinguish between different categories of targets.
[0208] After obtaining the target image, the target image is sent to the model, and the image processing network constructed by the 5-layer hierarchical iterative connection of the SPStem module, the SEDFF module, and the ADown module in the Backbone is used. The P1 to P5 layers of the image processing network each output corresponding neutral feature maps. The comprehensive feature maps output by the P3, P4, and P5 layers are selected, among which the feature map of the P3 layer has a relatively high resolution and is used to detect small-scale targets. The P4 layer feature map has a moderate resolution and is used to detect medium-scale targets. The P5 layer feature map has a low resolution and is used to detect large-scale targets.
[0209] The target is identified through the different scale feature maps corresponding to the size of the fruit target, and the identification result is output.
[0210] Compared with the related art, the application can realize high-precision detection of fruit targets: the SE-YOLO detection model reaches 93.6% and 67.3% in mAP@0.50 and mAP@0.50:0.95 indexes, respectively, which is better than existing high-precision detection methods such as RT-DETR, S-YOLO and Hyper-YOLO. The application can ensure lightweight deployment: only about 4.3% of the RT-DETR calculation cost is used, the parameter amount is 2.050M, the FLOPs is 5.6G, and the model can run efficiently on edge devices. It has good edge deployment capability: after INT8 quantization, it is deployed on the OrangePi Plus5 development board to realize real-time performance of 18.7FPS, while maintaining 88.8% of mAP@0.50 and 61.6% of mAP@0.50:0.95. Strong generalization ability: the effectiveness of the model is verified on multiple public datasets, including 2022Shanxi Nonggu Tomato Town Dataset and Cherry Tomato Datasets, which can effectively detect occluded and partially visible fruit targets such as tomatoes and maturity-related detection tasks. On a typical tomato dataset, the experimental results show that it reaches 93.6% and 67.3% in mAP@0.50 and mAP@0.50:0.95 indexes, respectively. Compared with existing advanced detectors, the model still maintains excellent detection performance with a significant reduction in calculation cost. In addition, the deployment experiment on edge devices shows that the model is superior to existing models in real-time performance and detection accuracy, proving its feasibility in actual agricultural scenarios.
[0211] Please refer to Figure 8 In the embodiments of the application, an identity recognition device 800 is provided, which includes but is not limited to an image acquisition module 810, a first image processing module 830, a second image processing module 850, a third image processing module 870 and a target recognition module 890.
[0212] The image acquisition module 810 is configured to acquire a target image containing a fruit target.
[0213] The first image processing module 830 is configured to compress the spatial features of the target image to generate a first pooled feature map. The edge features of the target image are extracted based on the Sobel operator to generate a first edge feature map, and the first pooled feature map and the first edge feature map are image spliced to generate a comprehensive feature map.
[0214] The second image processing module 850 is configured to perform convolution operation on the integrated feature map to generate a first convolution feature map, extract edge features in the integrated feature map based on a Sobel operator to generate a second edge feature map, and perform image fusion on the integrated feature map, the first convolution feature map and the second edge feature map based on a residual structure to update the integrated feature map with a generated second fusion image.
[0215] The third image processing module 870 is configured to perform average pooling processing on the integrated feature map and perform image segmentation to obtain a first feature submap and a second feature submap. The first feature submap is subjected to maximum pooling processing to generate a second pooling feature map. The second feature submap is subjected to convolution operation to generate a third convolution feature map. The second pooling feature map and the third convolution feature map are subjected to image splicing to update the integrated feature map with a generated fourth fusion image.
[0216] The target recognition module 890 is configured to acquire a target recognition model trained in advance based on a loss function including an attention component, a distance penalty component and a moving average strategy component, and input the integrated feature map into the pre-trained target recognition model to perform target recognition and generate an identification result of a target prediction box indicating a fruit target in the integrated feature map.
[0217] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the indication of the arrow, these steps are not necessarily executed in sequence according to the indication of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other sequences. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be alternately executed with at least part of other steps or sub-steps or stages of other steps.
[0218] The above is only some embodiments of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, some improvements and refinements can be made, which should also be regarded as the protection scope of the present application.
Claims
1. A method for detecting fruit targets, characterized in that, include: A target image containing the fruit target is acquired, and the target image is subjected to a first branch processing to compress the spatial features of the target image and generate a first pooling feature map, wherein the first branch processing is max pooling or average pooling. The second branch processing is performed based on the Sobel operator to extract the edge features of the target image and generate a first edge feature map, wherein the second branch processing is an edge detection processing; The first pooling feature map and the first edge feature map are image-stitched together to generate a comprehensive feature map. The comprehensive feature map is input into a pre-trained target recognition model for target recognition, generating a recognition result that indicates the target prediction box of the fruit target in the comprehensive feature map.
2. The method as described in claim 1, characterized in that, The Sobel operator includes a horizontal edge convolution kernel and a vertical edge convolution kernel; The second branch processing based on the Sobel operator to extract edge features of the target image to generate a first edge feature map includes: The target image is convolved with a horizontal edge convolution kernel to extract the horizontal edge features of the target image and generate a horizontal edge feature map. The target image is convolved with a vertical edge convolution kernel to extract the vertical edge features of the target image and generate a vertical edge feature map. The horizontal edge feature map and the vertical edge feature map are fused to generate a first edge feature map.
3. The method as described in claim 1, characterized in that, The step of concatenating the first pooling feature map and the first edge feature map to generate a comprehensive feature map includes: Channel adjustment is performed on the first pooling feature map and the first edge feature map to adjust the first pooling feature map and the first edge feature map to a first set number of channels; Based on the first set number of channels, the first pooling feature map and the first edge feature map are concatenated along the channel dimension to generate a first fused image; The first fused image is convolved to adjust the number of channels of the first fused image to a second set number of channels, thereby generating a comprehensive feature map, wherein the second set number of channels is less than the first set number of channels.
4. The method as described in claim 1, characterized in that, After concatenating the first pooling feature map and the first edge feature map to generate a comprehensive feature map, the method further includes performing the following operations on the comprehensive feature map: Perform a convolution operation on the comprehensive feature map to generate a first convolutional feature map; The edge features in the comprehensive feature map are extracted based on the Sobel operator to generate a second edge feature map; Based on the residual structure, the comprehensive feature map, the first convolutional feature map, and the second edge feature map are fused to generate a second fused image to update the comprehensive feature map.
5. The method as described in claim 4, characterized in that, The step of fusing the comprehensive feature map, the first convolutional feature map, and the second edge feature map based on the residual structure, and updating the comprehensive feature map with the generated second fused image, includes: The first convolutional feature map and the second edge feature map are fused to generate a third fused image; The third fused image is convolved to generate a second convolutional feature map; Based on the residual structure, the integrated feature map and the second convolutional feature map are connected, and image fusion is performed to generate a second fused image; The synthesized feature map is updated with the generated second fused image.
6. The method as described in claim 1, characterized in that, After concatenating the first pooling feature map and the first edge feature map to generate a comprehensive feature map, the method further includes: The composite feature map is subjected to average pooling to generate an average pooled feature map; The average pooling feature map is divided into a first feature sub-map and a second feature sub-map whose image channel dimensions are different from the first feature sub-map. Max pooling is performed on the first feature sub-map to generate a second pooled feature map; The second feature sub-map is convolved to generate a third convolutional feature map; The second pooling feature map and the third convolutional feature map are concatenated to generate a fourth fused image, which is then used to update the comprehensive feature map.
7. The method as described in claim 1, characterized in that, The target recognition model is trained based on a loss function that includes an attention component, a distance penalty component, and a moving average strategy component. The training method for the target recognition model includes: A pre-trained model is obtained and a training set is selected to train the target detection model for target recognition. The training images in the training set include fruit targets and their corresponding ground truth bounding boxes. The attention component calculates the degree of overlap between the ground truth bounding box and the target prediction box generated by the target detection model during the training process, and generates an attention result. The distance penalty component calculates the relative distance difference between the ground truth bounding box and the target prediction box during the training process to generate a distance penalty result. The moving average strategy component calculates the loss intensity based on the attention result and the distance penalty result, and updates the model parameters of the target detection model based on the loss intensity until the target recognition training is completed.
8. The method as described in claim 7, characterized in that, The attention-based component calculates the overlap between the ground truth bounding boxes and the target prediction boxes generated by the object detection model during the training process, generating attention results, including: The overlap between the ground truth bounding boxes and the target prediction boxes generated by the target detection model is calculated to generate an overlap parameter. Obtain attention parameters and set the weight of the target prediction box based on the difference between the overlap parameter and the attention parameter; Attention results are generated based on the weights of the target prediction boxes.
9. The method as described in claim 7, characterized in that, The moving average-based strategy component calculates a loss intensity based on the attention result and the distance penalty result, and updates the model parameters of the target detection model based on the loss intensity until the target recognition training is completed, including: Obtain the training phase corresponding to each loss intensity, wherein the training phase is pre-divided based on the model performance training requirements of the target recognition training; The model performance trend of the target recognition model is calculated based on the attention result, distance penalty result, and moving average policy component to determine the training stage of the target recognition training. The target recognition model parameters are updated based on the loss intensity corresponding to the training phase until the target recognition training is completed.
10. The method according to any one of claims 1-6, characterized in that, The step of inputting the comprehensive feature map into a pre-trained target recognition model for target recognition, and generating a recognition result indicating the target prediction box of the fruit target in the comprehensive feature map, includes: The comprehensive feature map is input into a hierarchically iteratively connected image processing network for image processing. In the image processing network, the comprehensive feature map output from the previous level is used as the target image for first branch processing and second branch processing, and the corresponding comprehensive feature map is output. Determine the image scale of the comprehensive feature map output by each layer of the image processing network during image processing; The image scale for target recognition is determined based on the size of the fruit target in the target image; The comprehensive feature map of the corresponding image scale is input into the target recognition model for target recognition, and the recognition result is generated.
11. A target detection device, characterized in that, include: The image acquisition module is used to acquire target images containing fruit objects; The first image processing module is used to compress the spatial features of the target image to generate a first pooling feature map; extract the edge features of the target image based on the Sobel operator to generate a first edge feature map; and stitch the first pooling feature map and the first edge feature map together to generate a comprehensive feature map. The second image processing module is used to perform a convolution operation on the comprehensive feature map to generate a first convolution feature map, extract edge features from the comprehensive feature map based on the Sobel operator to generate a second edge feature map, and perform image fusion of the comprehensive feature map, the first convolution feature map and the second edge feature map based on the residual structure to update the comprehensive feature map with the generated second fused image. The third image processing module is used to perform average pooling on the comprehensive feature map and perform image segmentation to obtain a first feature sub-map and a second feature sub-map; perform max pooling on the first feature sub-map to generate a second pooled feature map; perform convolution operation on the second feature sub-map to generate a third convolutional feature map; and stitch the second pooled feature map and the third convolutional feature map together to update the comprehensive feature map with the generated fourth fused image. The target recognition module is used to obtain a target recognition model that has been pre-trained based on a loss function including an attention component, a distance penalty component, and a moving average strategy component, and input the comprehensive feature map into the pre-trained target recognition model for target recognition, generating a target prediction box indicating the fruit target in the comprehensive feature map.