Facility crop fruit detection method and device based on multi-source heterogeneous image fusion
Through multi-source heterogeneous image fusion and generation of adversarial network reconstruction images, the problem of low accuracy in plant crop fruit detection under complex background and occlusion conditions is solved, and high-precision fruit detection is achieved.
Patent Information
- Application Number
- CN202510716184.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
AI Technical Summary
Under the conditions of complex background and interference factors with occlusion, the accuracy and recall of facility crop fruit detection based on visible light images are low, making it difficult to meet the demand for high precision and high reliability in precision agriculture.
Using a method based on multi-source heterogeneous image fusion, visible light images and infrared images are obtained, image reconstruction is carried out through image fusion and feature fusion, and fruit detection is performed on the reconstructed image.
It improves the accuracy and recall of plant crop fruit detection, and can achieve accurate detection under complex backgrounds and occlusion conditions to meet the needs of precision agriculture.
Smart Images

Figure CN120220142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly to a method and device for detecting fruits of protected crops based on multi-source heterogeneous image fusion. Background Art
[0002] In the process of the intelligent development of agriculture, the detection of fruits of protected crops, as a key link in precision agriculture, plays a crucial role in aspects such as yield prediction and growth trend monitoring. Most existing studies directly detect the visible light images of protected crops by means of an object detection network improved based on YOLOv8. Although in a relatively simple environment, effective recognition of the fruits of protected crops can be achieved. However, the growth environment of the plants of actual protected crops may be extremely complex, and there may be a large number of interference factors in the background of the visible light images, such as weeds, planting equipment brackets, etc. At the same time, the colors of some crop fruits are similar to those of the leaves, and some fruits are blocked by the leaves and the branches of the protected crops. Under these conditions of complex backgrounds and occlusion interference factors, the accuracy and recall rate of the detection of fruits of protected crops based on visible light images are low, and it is difficult to meet the requirements of precision agriculture for high-precision and high-reliability detection of fruits of protected crops. Summary of the Invention
[0003] The present invention provides a method and device for detecting fruits of protected crops based on multi-source heterogeneous image fusion to solve the defect of low accuracy and recall rate in the detection of fruits of protected crops under the conditions of complex backgrounds and occlusion interference factors.
[0004] The present invention provides a method for detecting fruits of protected crops based on multi-source heterogeneous image fusion, including: Obtaining multi-source heterogeneous images of the fruits of the protected crops to be detected, where the multi-source heterogeneous images include visible light images and infrared images; Performing image fusion on the multi-source heterogeneous images to obtain a true fusion image; Extracting features from the visible light image and the infrared image respectively, and fusing the extracted visible light image features and infrared image features to obtain fused features; Inputting the true fusion image and the fused features into a generative adversarial network for generation to obtain a reconstructed image of the fruits of the protected crops to be detected; Performing fruit detection on the reconstructed image to obtain the detection result of the fruits of the protected crops to be detected.
[0005] In some embodiments, after obtaining the multi-source heterogeneous images of the fruits of the protected crops to be detected, the method further includes: Determining the fruit coordinates of the fruits of the protected crops to be detected in the fruit annotation boxes of the multi-source heterogeneous images; Construct a fruit feature region based on the boundary values of the fruit coordinates; Determine the pixel range including the fruit feature region in the multi-source heterogeneous image; Crop the image region outside the pixel range in the multi-source heterogeneous image.
[0006] In some embodiments, the step of respectively performing feature extraction on the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fused features includes: Call a convolutional neural network to respectively extract the visible light image features of the visible light image and the infrared image features of the infrared image; According to the preset visible light feature weight and infrared feature weight, respectively weight the visible light image features and the infrared image features, and sum the weighted image features to obtain fused features, where the sum of the values of the visible light feature weight and the infrared feature weight is 1.
[0007] In some embodiments, the step of inputting the real fused image and the fused features into a generative adversarial network for generation to obtain a reconstructed image of the crop fruit of the facility to be detected includes: Call the generator in the generative adversarial network to generate the fused features to obtain a generated image; Call the discriminator in the generative adversarial network to discriminate between the generated image and the real fused image; When the discriminator recognizes the generated image as the real fused image, use the generated image as the reconstructed image of the crop fruit of the facility to be detected.
[0008] In some embodiments, the training process of the generative adversarial network includes: Construct a real fused image sample; Call the generator of the generative adversarial network to generate a random noise vector to obtain a first generated fused image; Input the first generated fused image and the real fused image sample into the discriminator of the generative adversarial network to obtain corresponding first generated discrimination results and real discrimination results; Construct a first loss function according to the first generated discrimination result and the real discrimination result, and perform first backpropagation in the generative adversarial network through the first loss function to update the parameters of the discriminator, where during the first backpropagation, the parameters of the generator remain fixed; Call the generator to generate a random noise vector to obtain a second generated fused image; Input the second generated fusion image into the discriminator with updated parameters to obtain the corresponding second generated discrimination result; Construct a second loss function according to the second generated discrimination result, and perform second backpropagation in the generative adversarial network through the second loss function to update the parameters of the generator. Wherein, during the second backpropagation, the parameters of the discriminator with updated parameters remain fixed.
[0009] In some embodiments, the fruit detection for the reconstructed image to obtain the detection result of the fruit of the facility crop to be detected includes: Call the YOLOv11 model to perform fruit detection on the reconstructed image to obtain the detection result of the fruit of the facility crop to be detected; Among them, the training process of the YOLOv11 model includes: Obtain the reconstructed image samples of the facility crop fruits, and the reconstructed image samples carry the labeled fruit region labels; Perform scale scaling on the reconstructed image samples to obtain multi-scale image samples; Input the reconstructed image samples and the multi-scale image samples into the YOLOv11 model for forward propagation to obtain the fruit prediction region; Construct a cross-entropy loss function according to the fruit region label and the fruit prediction region; Perform backpropagation in the YOLOv11 model through the cross-entropy loss function to update the parameters of the YOLOv11 model.
[0010] The present invention also provides a facility crop fruit detection device based on multi-source heterogeneous image fusion, including: An acquisition module, configured to acquire multi-source heterogeneous images of the fruit of the facility crop to be detected, and the multi-source heterogeneous images include visible light images and infrared images; A fusion module, configured to perform image fusion on the multi-source heterogeneous images to obtain a real fusion image; An extraction module, configured to respectively extract features from the visible light image and the infrared image, and fuse the extracted visible light image features and infrared image features to obtain fusion features; A reconstruction module, configured to input the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the fruit of the facility crop to be detected; A detection module, configured to perform fruit detection on the reconstructed image to obtain the detection result of the fruit of the facility crop to be detected.
[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for detecting fruits of protected crops based on multi-source heterogeneous image fusion as described in any one of the above is implemented.
[0012] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting fruits of protected crops based on multi-source heterogeneous image fusion as described in any one of the above is implemented.
[0013] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for detecting fruits of protected crops based on multi-source heterogeneous image fusion as described in any one of the above is implemented.
[0014] The method and device for detecting fruits of protected crops based on multi-source heterogeneous image fusion provided by the present invention respectively perform image fusion and feature fusion on two heterogeneous images, namely, visible light images and infrared images, of the fruits of protected crops to be detected, and use a generative adversarial network to achieve the fusion and reconstruction of visible light images and infrared images, ensuring the accuracy of image reconstruction and better meeting the actual detection requirements. Visible light images can capture reflected light information and have a high ability to distinguish texture details, while infrared images capture thermal radiation information and can, to a certain extent, eliminate the occlusion effect of leaves on fruits. Finally, the reconstructed images are used for fruit detection, which can accurately detect the fruits of protected crops under the conditions of complex backgrounds and interference factors such as occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic flowchart of the method for detecting fruits of protected crops based on multi-source heterogeneous image fusion provided by the present invention.
[0017] Figure 2 is a schematic framework principle diagram of the method for detecting fruits of protected crops based on multi-source heterogeneous image fusion provided by the present invention.
[0018] Figure 3 is a schematic diagram of the model architecture of the generative adversarial network provided by the present invention.
[0019] Figure 4It is a schematic structural diagram of the facility crop fruit detection device based on multi-source heterogeneous image fusion provided by the present invention.
[0020] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0021] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] The method and device for detecting facility crop fruits based on multi-source heterogeneous image fusion of the present invention will be described below with reference to the accompanying drawings. Figure 1 It is a schematic flow diagram of the method for detecting facility crop fruits based on multi-source heterogeneous image fusion provided by the present invention. As Figure 1 shown, the method includes the following steps 101 to 105.
[0023] Step 101: Obtain multi-source heterogeneous images of the facility crop fruits to be detected.
[0024] As Figure 2 shown, in multi-source data collection and preprocessing, multi-source data collection is first performed and then preprocessing is executed. During the process of planting the facility crop to be detected, image sensors can be installed at fixed positions in the planting area. For example, digital cameras or infrared cameras, etc. Then, the image sensor is used to collect the fruit image data during the growth process of the facility crop to be detected as multi-source data. The facility crop is generally an agricultural crop, such as tomatoes, etc. The multi-source data collected is multi-source heterogeneous images, specifically including visible light images and infrared images. The directly collected multi-source heterogeneous images generally need to be preprocessed. The preprocessing process includes image cleaning and image annotation. Image cleaning is to delete some images with unsatisfactory resolution or large influence of light factors. For example, some images without the facility crop fruits to be detected also need to be cleaned. And image annotation is to mark the positions of the fruits in the images. For the multi-source heterogeneous images after image cleaning, the positions of the fruits therein can be pre-annotated by manual annotation, and the fruit regions in the images can be framed by means of fruit annotation frames (such as rectangular frames).
[0025] In some embodiments, as Figure 2 shown, after obtaining the multi-source heterogeneous images of the facility crop fruits to be detected, the next step is to perform precise construction of the fruit feature regions, which is specifically divided into three parts: extracting fruit coordinates, constructing fruit feature regions, and cropping the feature regions. The following is a specific description.
[0026] First, extract the fruit coordinates. In the fruit annotation boxes of the multi-source heterogeneous images, determine the fruit coordinates of the facility crop fruits to be detected. Since the multi-source heterogeneous images after image cleaning have had the fruit regions framed through manual annotation, it is possible to directly extract and record the fruit coordinates of the facility crop fruits to be detected in the fruit annotation boxes of the multi-source heterogeneous images using annotation tools such as LabelImg.
[0027] To ensure that more fruits are included in the fruit feature region, when constructing the fruit feature region here, construct the fruit feature region according to the boundary values of the fruit coordinates. By extracting the boundary values from all the fruit coordinates, denoted as , , , . These boundary values represent the maximum and minimum coordinate values of the facility crop fruits in the horizontal x-axis direction and the maximum and minimum coordinate values in the vertical y-axis direction. In this way, the fruit feature region can be constructed according to the boundary values. Define the length of the fruit feature region as , and the width as .
[0028] Finally, for feature region cropping, first, it is necessary to determine the pixel range including the fruit feature region in the multi-source heterogeneous image. Since there may be multiple facility crops (fruits) in the multi-source heterogeneous image, there may be multiple fruit feature regions. Therefore, determine the pixel range of the fruits from the multi-source heterogeneous image here. This pixel range must ensure that all the fruit feature regions are included, and this pixel range can be a rectangular region range. Since there are no facility crops in the image regions outside the pixel range, there is no need to perform object detection. Therefore, next, crop the image regions outside the pixel range in the multi-source heterogeneous image and only retain the pixel range including the fruit feature region.
[0029] It should be noted that the multi-source heterogeneous images include visible light images and infrared images. Both of these heterogeneous images are processed in the same way as described above during cleaning and preprocessing, that is, the three processes of extracting fruit coordinates, constructing fruit feature regions, and feature region cropping are performed respectively.
[0030] In the embodiment of the present invention, before performing fruit detection, cleaning and preprocessing are performed on the obtained multi-source heterogeneous images, which can effectively improve the quality of the image data. And through the annotation of the fruit feature region and the cropping process of the non-fruit feature region, not only can the image information be focused, making subsequent image reconstruction and fruit detection more accurate, but also the information irrelevant to the fruit features in the image can be reduced, the image size can be reduced, and the computational amount of the subsequent detection model can be reduced.
[0031] Step 102: Perform image fusion on the multi-source heterogeneous images to obtain a real fused image.
[0032] After the multi-source heterogeneous images are cleaned and preprocessed, the visible light images and infrared images included in the multi-source heterogeneous images are fused to obtain a real fused image. The fusion method is pixel-level fusion, where two pixels to be fused are added pixel by pixel. It can also be achieved by using image fusion algorithms such as multi-scale transformation or by using a deep learning model.
[0033] Step 103: Extract features from the visible light image and the infrared image respectively, and fuse the extracted visible light image features and infrared image features to obtain fused features.
[0034] The purpose of image fusion is to superimpose the feature information of two images. Therefore, in the embodiments of the present invention, a generative adversarial network (GAN) is used to achieve image reconstruction in the feature dimension. First, features are extracted from the visible light image and the infrared image respectively, and the extracted visible light image features and infrared image features are fused to obtain fused features. The extraction method can be implemented by using an image encoder such as a Transformer or a deep learning model such as a convolutional neural network. The visible light image features are specifically the visual feature information such as texture, color, and shape in the image, denoted as for characterizing the texture, color, shape, etc. of the crop fruits of the facility to be detected, while the infrared image features are specifically the thermal radiation feature information of the crop fruits of the facility to be detected, denoted as . The finally fused features can express the information of both of these features at the same time.
[0035] As Figure 2 shown, in the process of multi-source image fusion based on GAN, first, visible light image features and infrared image features are extracted respectively, and then weighted fusion of the features is performed for the final image reconstruction.
[0036] Specifically, first, a convolutional neural network is called to extract the visible light image features of the visible light image and the infrared image features of the infrared image respectively. That is to say, the extraction method of both features can be implemented by using a convolutional neural network. Next, weighted fusion is performed. According to the preset visible light feature weight and infrared feature weight, the visible light image features and the infrared image features are weighted respectively, and the weighted image features are summed to obtain fused features, denoted as , as shown in the following formula: (1) In the above formula (1), Represents the preset visible light feature weight, Represents the preset infrared feature weight, Represents a multiplication operation. Here, the sum of the values of the visible light feature weight and the infrared feature weight is 1. These two weights can be preset through experimental verification or prior knowledge, and the values range from 0 to 1.
[0037] In the embodiments of the present invention, the weighted fusion of visible light image features and infrared image features is achieved through the weight distribution results verified by experiments, which can ensure that the fused features can better express visual feature information and thermal radiation feature information.
[0038] Step 104: Input the real fused image and the fused features into a generative adversarial network for generation to obtain a reconstructed image of the crop fruits of the facility to be detected.
[0039] In the final image reconstruction stage, the real fused image and the fused features are input into a generative adversarial network for generation to obtain a reconstructed image of the crop fruits of the facility to be detected. The generative adversarial network includes a generator G and a discriminator D. The generator can reconstruct an image according to the input fused features To reconstruct the image, the generator G maps the low-dimensional fused features back to the high-dimensional image space through a series of deconvolution operations or transposed convolution operations, gradually restoring the details and structure of the image to generate a reconstructed image. At the same time, the discriminator D judges the generated reconstructed image, compares the gap between the reconstructed image and the input real fused image, and feeds it back to the generator to prompt the generator to continuously improve, making the reconstructed image more realistic and meeting the requirements.
[0040] The following specifically describes the generation process in the generative adversarial network. First, call the generator in the generative adversarial network to generate the fused features to obtain a generated image. As Figure 3 shown, the generative adversarial network here is specifically a deep convolutional generative adversarial network. The generator G is a five-layer convolutional neural network. The first and second layers use a 5×5 convolutional kernel conv as a filter, the third and fourth layers use a 3×3 convolutional kernel conv as a filter, and the last layer uses a 1×1 convolutional kernel conv as a filter. The stride in each layer is set to 1, and there is no padding operation in the convolution. The first four layers adopt batch normalization BatchNorm, which can make the model more stable and help the gradient to effectively backpropagate to each layer. For the activation function, the Leaky ReLU activation function is used in the first four layers, and the tanh activation function is used in the last layer.
[0041] Through the 5-layer convolutional neural network in the generator G, the fused features can be mapped back to the high-dimensional image space, gradually restoring the details and structure of the image, and finally outputting the generated image.
[0042] Next, the discriminator D in the generative adversarial network is called to discriminate the generated image and the real fused image. The discriminator D is used to judge the gap between the generated image and the real fused image and score it. Its judgment process is specifically to perform classification, and the score is used to determine whether the generated image and the real fused image belong to the same class. As Figure 3 shown, the discriminator D is also a five-layer neural network. From the first layer to the fourth layer, it is a convolutional neural network. The 3×3 convolutional kernel is used as the filter, the stride is set to 2, and there is no padding operation in the convolution. The activation function is the Leaky Relu function. Among them, batch normalization (BatchNorm) is adopted in the second layer and the fourth layer. The last layer is a linear layer (Linear) for classification and outputting the final discrimination result.
[0043] After the discriminator D outputs the discrimination result, when the discrimination result indicates that the generated image and the real fused image are not of the same class, that is, the gap between the generated image and the real fused image is large, the discrimination result needs to be fed back to the generator G for regeneration, so that the generated image output by the generator G gets closer and closer to the real fused image.
[0044] When the discriminator D recognizes the generated image as the real fused image, the generated image is used as the reconstructed image of the crop fruit of the facility to be detected. That is to say, the generated image output by the generator G has approached the real fused image, resulting in the discriminator D being unable to distinguish, and recognizing the generated image as the real fused image and belonging to the same class, indicating that the image reconstruction is completed. At this time, the generated image output by the generator G can be used as the reconstructed image of the crop fruit of the facility to be detected.
[0045] In the embodiment of the present invention, taking the real fused image as the target of image reconstruction and using the generative adversarial network to perform image reconstruction can generate a reconstructed image that simultaneously contains thermal radiation information and texture details under complex backgrounds or partial occlusions, realizing the feature information fusion of visible light images and infrared images, and having obvious advantages in enhancing target information and suppressing interference information.
[0046] The training process of the generative adversarial network is described below. In the embodiment of the present invention, the training process of the generative adversarial network is divided into two stages, and the two stages are alternately trained. In the first stage, the parameters of the generator G are fixed, and only the parameters of the discriminator D are trained. In the second stage, after the parameters of the discriminator D are trained and fixed, only the parameters of the generator G are trained.
[0047] First, construct real fused image samples. For example, a part of the real fused images obtained in step 102 above can be randomly sampled as real fused image samples, denoted as .
[0048] During the first-stage training process, the generator G of the generative adversarial network is called to generate a first generated fusion image from a random noise vector. The random noise vector z generally follows a Gaussian distribution or a uniform distribution. The random noise vector is input into the generator and, through the mapping process of the five-layer convolutional neural network of the generator G, the corresponding first generated fusion image is obtained, denoted as .
[0049] Then, the first generated fusion image and the real fusion image samples are input into the discriminator of the generative adversarial network to obtain the corresponding first generated discrimination result and the real discrimination result. The discriminator D can output the corresponding first generated discrimination result for the first generated fusion image , and for the real fusion image samples, output the corresponding real discrimination result .
[0050] Next, according to the first generated discrimination result and the real discrimination result , a first loss function is constructed, and the formula is as follows: (2) Finally, through the first loss function , the first backpropagation is performed in the generative adversarial network to update the parameters of the discriminator D. That is, during backpropagation, the parameters of the generator G are fixed, and only the parameters of the discriminator D are trained. During backpropagation, the gradient of the first loss function with respect to the parameters of the discriminator is calculated, and then the Adam optimizer is used to update the parameters of the discriminator so that the discriminator D can better distinguish between real images and generated images. The update formula is: (3) In the above formula (3), is the learning rate, represents the parameters of the discriminator, denotes the corresponding gradient, and
[0051] In the second stage, the generator G is still called to generate a second generated fusion image from the random noise vector z, denoted as, and then the second generated fusion image is directly input into the discriminator D with updated parameters to obtain the corresponding second generated discrimination result, denoted as .
[0052] Next, according to the second generated discrimination result Construct the second loss function , which is expressed by the formula as follows: (4) Finally, perform the second backpropagation in the generative adversarial network through the second loss function to update the parameters of the generator. Here, calculate the gradient of the second loss function with respect to the parameters of the generator G, and then use the Adam optimizer to update the parameters of the generator G, so that the images generated by the generator G can better approximate the real fusion image samples. Its update formula is: (5) In the above formula (5), is the learning rate, represents the parameters of the generator G, represents the corresponding gradient, represents the second loss function.
[0053] During each iteration in the training process, repeat the processes of the above two stages, alternately train the discriminator and the generator, and update the corresponding parameters through the first loss function and the second loss function . Stop the iteration and end the training process until the quality of the images generated by the generator reaches a satisfactory effect, or the discriminator and the generator reach an equilibrium state (that is, the first loss function and the second loss function start to converge).
[0054] In the embodiment of the present invention, the generative adversarial network is trained through a two-stage alternating training process, so that the images generated by the generator can be closer to the real fusion images, effectively retain the characteristic information of visible light images and infrared images, and ensure the effective detection of facility crop fruits under complex backgrounds and occlusion conditions.
[0055] Step 105: Perform fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected.
[0056] Finally, perform fruit detection on the reconstructed image generated by the generative adversarial network. Here, models such as the YOLO object detection network can be used to detect the facility crop fruit to be detected. Finally, the model can mark the real area or fruit contour of the facility crop fruit to be detected in the multi-source heterogeneous images as the detection result of the facility crop fruit to be detected.
[0057] Such as Figure 2As shown in the figure, in the crop fruit detection stage, the embodiment of the present invention constructs a crop fruit detection model based on YOLOv11, and performs fruit detection on the reconstructed image through the YOLOv11 model to obtain the detection result of the facility crop fruit to be detected.
[0058] Since the YOLOv11 model performs fruit detection on the reconstructed image, the embodiment of the present invention trains the YOLOv11 model by constructing a reconstructed image sample before performing fruit detection. The training process of the YOLOv11 model is introduced below.
[0059] First, obtain the reconstructed image samples of the facility crop fruits. Here, the above steps 101 and 102 can be executed. Obtain the visible light image and the infrared image through step 101, and then execute step 102 to obtain the corresponding real fusion image as the reconstructed image sample, that is, the image samples of the facility crop fruits covering different growth stages (green fruit stage, color-changing stage, mature stage), different lighting conditions (strong light on sunny days, weak light on cloudy days), and different occlusion degrees (partially occluded by branches and leaves, completely exposed).
[0060] Here, the reconstructed image samples carry labeled fruit region tags. Similarly, here, the fruit regions of each reconstructed image sample are also labeled manually, and each fruit region is labeled with a rectangular box. These rectangular boxes are the fruit region tags carried in the reconstructed image samples.
[0061] In order to enable the YOLOv11 model to detect facility crop fruits of various sizes, a multi-scale training strategy is adopted when training the YOLOv11 model, and the reconstructed image samples are scaled to obtain multi-scale image samples. Here, the reconstructed image samples are respectively scaled by 0.5 times and 2 times to obtain images of different sizes.
[0062] Then, the reconstructed image samples and the multi-scale image samples are input into the YOLOv11 object detection model for forward propagation to obtain the fruit prediction regions. During the forward propagation process, through processes such as feature extraction and prediction classification, the facility crop fruits in the sample images can be recognized, and the fruit prediction regions are determined. Then, the fruit prediction regions are compared with the fruit region tags, and a cross-entropy loss function is constructed based on the fruit region tags and the fruit prediction regions to label the difference between the model prediction and the true label.
[0063] Finally, backpropagation is performed in the YOLOv11 model through the cross-entropy loss function to update the parameters of the YOLOv11 model. During the backpropagation process, the gradients of the model are calculated, and the gradients are optimized to update the parameters of the model. Before training, 70% of the reconstructed image samples are used as the training set, and the remaining are used as the test set. Iterations are performed during the training process. When the number of iterations after a certain iteration reaches the total number of preset iterations or the cross-entropy loss function starts to converge, the iteration is stopped, and the training process ends. The trained YOLOv11 model can be directly used to detect fruits in the reconstructed image obtained in step 104 to obtain the detection results of the fruits of the facility crops to be detected.
[0064] Thus, in the embodiment of the present invention, the YOLOv11 model is trained through the real fusion image of the facility crop fruits, so that the model has both the ability to distinguish texture details of visible light images and the characteristic of obtaining thermal radiation information by penetrating occlusion of infrared images, effectively solving the problem of fruit detection under complex backgrounds and occlusion conditions and making up for the deficiencies of traditional single-image detection. Moreover, through the multi-scale training strategy, the model can also have the detection ability for facility crop fruits of different scales, realizing accurate and effective detection of facility crop fruits.
[0065] The following describes the facility crop fruit detection device based on multi-source heterogeneous image fusion provided by the present invention. The facility crop fruit detection device based on multi-source heterogeneous image fusion described below can be correspondingly referred to the facility crop fruit detection method based on multi-source heterogeneous image fusion described above.
[0066] As Figure 4 shown, the facility crop fruit detection device based on multi-source heterogeneous image fusion specifically includes: an acquisition module 401, a fusion module 402, an extraction module 403, a reconstruction module 404, and a detection module 405. Among them, the acquisition module 401 is used to acquire multi-source heterogeneous images of the facility crop fruits to be detected, and the multi-source heterogeneous images include visible light images and infrared images; the fusion module 402 is used to perform image fusion on the multi-source heterogeneous images to obtain a real fusion image; the extraction module 403 is used to respectively extract features from the visible light image and the infrared image, and fuse the extracted visible light image features and infrared image features to obtain fusion features; the reconstruction module 404 is used to input the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the facility crop fruits to be detected; the detection module 405 is used to perform fruit detection on the reconstructed image to obtain the detection results of the facility crop fruits to be detected.
[0067] It should be noted that the beneficial effects of the facility crop fruit detection device based on multi-source heterogeneous image fusion here correspond to those of the facility crop fruit detection method based on multi-source heterogeneous image fusion in the above text. Therefore, the beneficial effects of the facility crop fruit detection device based on multi-source heterogeneous image fusion will not be elaborated here.
[0068] Figure 5 An example of a schematic physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the facility crop fruit detection method based on multi-source heterogeneous image fusion. The method includes: acquiring multi-source heterogeneous images of the facility crop fruit to be detected, where the multi-source heterogeneous images include visible light images and infrared images; performing image fusion on the multi-source heterogeneous images to obtain a real fusion image; respectively extracting features from the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fusion features; inputting the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the facility crop fruit to be detected; performing fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected.
[0069] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0070] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the facility crop fruit detection method based on multi-source heterogeneous image fusion provided by the above-mentioned various methods. The method includes: obtaining multi-source heterogeneous images of the facility crop fruit to be detected, where the multi-source heterogeneous images include visible light images and infrared images; performing image fusion on the multi-source heterogeneous images to obtain a real fusion image; respectively extracting features from the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fusion features; inputting the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the facility crop fruit to be detected; and performing fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected.
[0071] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the facility crop fruit detection method based on multi-source heterogeneous image fusion provided by the above-mentioned various methods. The method includes: obtaining multi-source heterogeneous images of the facility crop fruit to be detected, where the multi-source heterogeneous images include visible light images and infrared images; performing image fusion on the multi-source heterogeneous images to obtain a real fusion image; respectively extracting features from the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fusion features; inputting the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the facility crop fruit to be detected; and performing fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected.
[0072] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0073] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting facility crop fruits based on multi-source heterogeneous image fusion, characterized in that, Including: Obtain multi-source heterogeneous images of the crop fruits of the facility to be detected, where the multi-source heterogeneous images include visible light images and infrared images; Perform image fusion on the multi-source heterogeneous images to obtain a real fusion image; Extract features from the visible light image and the infrared image respectively, and fuse the extracted visible light image features and infrared image features to obtain fusion features; Input the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the crop fruits of the facility to be detected; Perform fruit detection on the reconstructed image to obtain the detection result of the crop fruits of the facility to be detected.
2. The facility crop fruit detection method based on multi-source heterogeneous image fusion according to claim 1, characterized in that, After obtaining the multi-source heterogeneous images of the crop fruits of the facility to be detected, the method further includes: Determine the fruit coordinates of the crop fruits of the facility to be detected in the fruit annotation box of the multi-source heterogeneous images; Construct a fruit feature region according to the boundary values of the fruit coordinates; Determine the pixel range including the fruit feature region in the multi-source heterogeneous images; Crop the image region other than the pixel range in the multi-source heterogeneous images.
3. The facility crop fruit detection method based on multi-source heterogeneous image fusion according to claim 1, characterized in that, The step of respectively extracting features from the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fusion features includes: Call a convolutional neural network to extract the visible light image features of the visible light image and the infrared image features of the infrared image respectively; According to the preset visible light feature weight and infrared feature weight, weight the visible light image features and the infrared image features respectively, and sum the weighted image features to obtain fusion features, where the sum of the values of the visible light feature weight and the infrared feature weight is 1.
4. The facility crop fruit detection method based on multi-source heterogeneous image fusion according to claim 1, characterized in that, The step of inputting the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the crop fruits of the facility to be detected includes: Call the generator in the generative adversarial network to generate the fusion features to obtain a generated image; Call the discriminator in the generative adversarial network to discriminate the generated image and the real fusion image; When the discriminator recognizes the generated image as the real fusion image, use the generated image as the reconstructed image of the crop fruits of the facility to be detected.
5. The facility crop fruit detection method based on multi-source heterogeneous image fusion according to claim 1, wherein The training process of the generative adversarial network includes: Construct real fusion image samples; Call the generator of the generative adversarial network to generate a random noise vector to obtain a first generated fusion image; Input the first generated fusion image and the real fusion image samples into the discriminator of the generative adversarial network to obtain corresponding first generated discrimination results and real discrimination results; Construct a first loss function according to the first generated discrimination result and the real discrimination result, and perform first backpropagation in the generative adversarial network through the first loss function to update the parameters of the discriminator, where during the first backpropagation, the parameters of the generator remain fixed; Call the generator to generate a random noise vector to obtain a second generated fusion image; Input the second generated fusion image into the discriminator with updated parameters to obtain the corresponding second generated discrimination result; Construct a second loss function according to the second generated discrimination result, and perform second backpropagation in the generative adversarial network through the second loss function to update the parameters of the generator. Wherein, during the second backpropagation, the parameters of the discriminator with updated parameters remain fixed.
6. The facility crop fruit detection method based on multi-source heterogeneous image fusion according to claim 1, characterized in that Performing fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected, including: Call the YOLOv11 model to perform fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected; Among them, the training process of the YOLOv11 model includes: Obtain a reconstructed image sample of the facility crop fruit, and the reconstructed image sample carries a labeled fruit region label; Perform scale scaling on the reconstructed image sample to obtain a multi-scale image sample; Input the reconstructed image sample and the multi-scale image sample into the YOLOv11 model for forward propagation to obtain a fruit prediction region; Construct a cross-entropy loss function according to the fruit region label and the fruit prediction region; Perform backpropagation in the YOLOv11 model through the cross-entropy loss function to update the parameters of the YOLOv11 model.
7. An apparatus for detecting fruits of protected crops based on multi-source heterogeneous image fusion, characterized in that, Including: An acquisition module for acquiring multi-source heterogeneous images of the facility crop fruit to be detected, where the multi-source heterogeneous images include visible light images and infrared images; A fusion module for fusing the multi-source heterogeneous images to obtain a real fusion image; An extraction module for respectively extracting features from the visible light image and the infrared image, and fusing the extracted visible light image features and infrared image features to obtain fusion features; A reconstruction module for inputting the real fusion image and the fusion features into a generative adversarial network for generation to obtain a reconstructed image of the facility crop fruit to be detected; A detection module for performing fruit detection on the reconstructed image to obtain the detection result of the facility crop fruit to be detected.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the facility crop fruit detection method based on multi-source heterogeneous image fusion according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the facility crop fruit detection method based on multi-source heterogeneous image fusion according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the facility crop fruit detection method based on multi-source heterogeneous image fusion according to any one of claims 1 to 6.
Citation Information
Patent Citations
Visible light and infrared image fusion method and system, storage medium and terminal
CN115100089A
Infrared and visible light image fusion method of multi-scale generative adversarial network
CN116863285A
Unmanned aerial vehicle intelligent wild animal monitoring method based on FPGA and improved YOLO
CN119206770A