Transplanting posture quality detection method based on multi-modal perception

By fusing RGB and multispectral images in farmland seedling attitude quality detection, and designing a multi-domain interactive prompt module and a cross-band attention fusion module, the problem of low detection accuracy in complex environments is solved, and high-precision seedling attitude quality detection is achieved.

CN120107798APending Publication Date: 2025-06-06HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510262341.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing farmland seedling attitude quality detection technology has poor adaptability in complex environments and low accuracy in drone farmland seedling attitude detection.

Method used

A method of transplanting posture quality detection based on multimodal perception is proposed. By fusing RGB images and multispectral images acquired by the UAV, a multi-domain interaction prompt module and a cross-band attention fusion module are designed to enhance the characteristic representation of seedlings, suppress background interference, and improve detection accuracy.

Benefits of technology

It significantly improves the detection accuracy of seedling targets and the recognition ability of seedlings in different postures under complex backgrounds, enhances the robustness and generalization capabilities of the model, and realizes high-precision detection of different postures such as floating seedlings, injured seedlings, and reasonable seedlings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107798A_ABST
    Figure CN120107798A_ABST
Patent Text Reader

Abstract

The invention relates to a transplanting posture quality detection method based on multi-modal perception, in particular to a transplanting posture quality detection method based on multi-modal perception. The invention aims to solve the problems that the existing farmland seedling attitude quality detection technology is poor in adaptability in a complex environment and the unmanned aerial vehicle farmland seedling attitude detection accuracy is low. The method comprises the following steps: collecting farmland RGB images and multispectral images in a seedling growth period; obtaining the preprocessed farmland RGB image and multispectral image in the seedling growing period; inputting the preprocessed farmland RGB image in the seedling growth period and the multispectral image into a deep learning model for training to obtain a trained deep learning model; based on the trained deep learning model, performing classification prediction on the farmland RGB image and the multispectral image of the to-be-detected seedling growing period, and predicting that seedlings in the farmland RGB image and the multispectral image of the to-be-detected seedling growing period are qualified seedlings, damaged seedlings or floating seedlings. The method is applied to the field of unmanned aerial vehicle farmland seedling attitude detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a rice transplanting posture quality detection method based on multi-modal perception. Background Art

[0002] With the in-depth development of the concept of precision agriculture and the increasing maturity of drone technology, the use of drone remote sensing technology to intelligently monitor the growth of crops has become an important part of modern agriculture. As a major food crop, the plant phenotype of rice during its growth process is directly related to the final yield and quality. In the rice production process, the transplanting link is crucial. The posture quality of the seedlings (such as floating seedlings, injured seedlings, and reasonable seedlings) directly affects the survival rate, tillering ability, and subsequent growth and development of rice. Therefore, rapid and accurate detection of the posture quality of farmland seedlings is of great significance for guiding agricultural production and improving management efficiency.

[0003] At present, manual inspection is the main way to detect the posture quality of farmland seedlings, but this method has disadvantages such as low efficiency, high labor intensity, and strong subjectivity, which makes it difficult to meet the needs of modern agricultural refined management. Although the traditional detection method based on ground observation has improved efficiency to a certain extent, it is difficult to achieve a rapid census of large areas of farmland due to the limitations of field of view and operation efficiency. UAV remote sensing technology, with its advantages of high mobility, wide coverage and low cost, provides a new solution for the automated and intelligent detection of the posture quality of seedlings in large areas of farmland. UAVs equipped with high-resolution visible light cameras can obtain RGB images of farmland, providing basic data for the identification of seedling targets. However, relying solely on RGB images for seedling posture quality detection faces many challenges. First, the farmland environment is highly complex, with many background interference factors, including soil, water and weeds. These factors have certain similarities with seedlings in color, texture and other characteristics, which can easily lead to false detection and missed detection in the target detection process; secondly, the apparent differences of seedlings in different postures on RGB images may not be significant. For example, slightly injured seedlings may be difficult to distinguish from healthy seedlings; in addition, a single RGB image lacks vegetation physiological and biochemical information, making it difficult to accurately judge the health status of seedlings.

[0004] In recent years, multispectral imaging technology has been widely used in the agricultural field. Multispectral cameras can obtain reflectance information of ground objects in multiple specific spectral bands, providing a richer data source for monitoring crop growth conditions. Seedlings with different postures and health conditions have different reflectance characteristics in different spectral bands. For example, healthy seedlings have a higher reflectance in the near-infrared band, while the reflectance of damaged seedlings will decrease. The spectral reflectance characteristics of floating seedlings will also change due to damaged roots and reduced water absorption capacity. Fusion of multispectral information with RGB information can more comprehensively describe the characteristics of seedlings and improve the accuracy and robustness of posture quality detection. However, how to effectively fuse multi-source heterogeneous RGB and multispectral data and fully explore the complementary information between different modal data is still a key issue that has not been fully studied in the field of transplanting quality detection.

[0005] Deep learning technology has made significant progress in the field of image recognition and target detection, providing a new technical means for the detection of farmland seedling posture quality. Target detection models based on deep learning, such as the YOLO (You Only Look Once) series, can achieve fast and accurate recognition of targets in images. However, there are still some shortcomings in directly applying general target detection models to the detection of farmland seedling posture quality. First, the seedling targets in farmland scenes are usually small in scale, densely distributed, and easily blocked by other objects, which puts higher requirements on the detection accuracy and robustness of the model; secondly, the differences in visual features of seedlings with different postures may be subtle, requiring the model to have stronger feature extraction and discrimination capabilities; therefore, it is necessary to combine the characteristics of farmland seedling posture quality detection and design an efficient and accurate deep learning model and algorithm to meet the actual needs of the task. Summary of the invention

[0006] The purpose of the present invention is to solve the problems that the existing farmland seedling posture quality detection technology has poor adaptability in complex environments and the accuracy of drone farmland seedling posture detection is low, and a rice transplanting posture quality detection method based on multimodal perception is proposed.

[0007] A method for detecting the quality of transplanting posture based on multimodal perception. The specific process is as follows:

[0008] Step S1, collecting RGB images and multispectral images of farmland during the seedling growth period;

[0009] Step S2, preprocessing the RGB image and multispectral image of the farmland during the seedling growth period to obtain the preprocessed RGB image and multispectral image of the farmland during the seedling growth period;

[0010] Step S3, constructing a deep learning model, inputting the pre-processed RGB images and multispectral images of the farmland during the seedling growth period into the deep learning model for training, and obtaining a trained deep learning model;

[0011] Step S4: classify and predict the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested based on the trained deep learning model, and predict the category of the seedlings in the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested as qualified seedlings, injured seedlings or floating seedlings.

[0012] The beneficial effects of the present invention are:

[0013] The present invention proposes a method for detecting the posture quality of farmland seedlings using a drone based on multimodal perception. The method integrates RGB images and multispectral images acquired by drones, fully utilizes the advantages of different modal data, and designs a multi-domain interactive prompt module and a cross-band attention fusion module to enhance the representation of seedling features, suppress background interference, and improve detection accuracy. At the same time, the present invention designs an online difficult example mining strategy to guide model training and improve the robustness and generalization ability of the model. The present invention aims to provide an efficient, accurate, and reliable drone farmland seedling posture quality detection solution to provide technical support for precision agricultural production.

[0014] In view of the shortcomings of existing farmland seedling posture quality detection technology in terms of adaptability to complex environments, multimodal information fusion and intelligent analysis, the present invention is committed to breaking through the bottleneck of existing technologies and constructing a set of drone farmland seedling posture quality accurate detection methods based on multimodal perception, aiming to achieve accurate identification of different postures such as floating seedlings, injured seedlings, and reasonable seedlings, and provide technical support for improving agricultural management and ensuring food security. The core goal of the present invention is to solve the following key technical problems:

[0015] The present invention needs to solve the problem of accurate identification of seedling targets in complex farmland backgrounds. The farmland environment is complex and changeable, the lighting conditions are uneven, and there are various interference factors such as soil, water, and weeds. These factors have certain similarities with the seedlings in color, texture and other characteristics, which brings challenges to the accurate identification of seedling targets. Especially from the perspective of drones, small-scale seedling targets are easily submerged by background noise, making it difficult for traditional image processing methods to effectively distinguish them. Therefore, how to design a detection method that can effectively suppress background interference and highlight the characteristics of seedling targets is the primary technical problem to be solved by the present invention.

[0016] The present invention needs to solve the problem of how to effectively fuse the multispectral image information of unmanned aerial vehicles to improve the accuracy of seedling posture quality detection. A single RGB image can only provide limited visible light information, and it is difficult to accurately distinguish seedlings of different postures. For example, the apparent characteristics of floating seedlings and injured seedlings on RGB images may be similar to those of healthy seedlings, making it difficult to effectively distinguish them. Multispectral images can provide rich physiological and biochemical information of vegetation. Seedlings of different postures have different reflectances in different spectral bands, which provides the possibility for more refined posture quality detection. However, there are differences between RGB images and multispectral images in data format, spatial resolution, etc. How to effectively fuse the data of these two modalities and fully tap the complementary information between them is one of the key technical problems to be solved by the present invention.

[0017] The present invention also needs to solve the problem of how to improve the model's fine-grained recognition ability for seedlings in different postures. The differences in apparent features between seedlings in different postures may be relatively subtle. For example, slightly injured seedlings and healthy seedlings may only have slight differences in color and shape. Traditional image classification or target detection methods may find it difficult to effectively capture these subtle differences, resulting in low recognition accuracy. Therefore, how to design a method that can effectively extract and distinguish the subtle features of seedlings in different postures is a key technical problem that the present invention needs to solve.

[0018] In summary, the present invention aims to provide a method for accurately detecting the posture quality of farmland seedlings by UAV based on multimodal perception. By fusing RGB and multispectral image information, the problem of accurate identification of seedlings under complex background interference is solved, the model's fine-grained discrimination ability for seedlings with different postures is improved, and the model's learning ability for difficult samples is enhanced, thereby achieving high-precision detection of different postures such as floating seedlings, injured seedlings, and reasonable seedlings, providing important technical support for refined agricultural management.

[0019] The present invention proposes a method for detecting the posture quality of farmland seedlings using drones based on multimodal perception. By effectively integrating the RGB images and multispectral image information collected by the drone platform and combining it with deep learning technology, the present invention achieves accurate and efficient evaluation of the posture quality of farmland seedlings. Compared with traditional manual inspection methods and detection methods that rely only on single-modal data, the present invention has achieved significant technical advantages in detection accuracy, robustness, and information fusion capabilities, providing strong technical support for precision agricultural production.

[0020] Specifically, the multi-domain interactive prompt module (MDIPM) designed in the present invention can effectively suppress the interference of complex farmland background and improve the model's attention to potential seedling areas. The module extracts image saliency information from multiple angles such as frequency domain, color space and learnable attention mechanism, guides the network to focus on seedling targets, reduces the false detection rate caused by background factors such as soil, water, and weeds, and significantly improves the detection rate and positioning accuracy of seedling targets. In addition, the cross-band attention fusion module (CBAF) constructed by the present invention realizes the effective fusion of RGB images and multispectral images, and makes full use of the rich physiological and biochemical characteristics of vegetation contained in multispectral information. The CBAF module enables RGB features to focus on specific spectral information related to different seedling postures, and enhances the model's recognition ability for abnormal postures such as floating seedlings and injured seedlings. The online difficult example mining strategy for seedlings designed by the present invention effectively avoids overfitting of the model on a large number of simple samples by dynamically screening out samples with high model prediction uncertainty for key learning during the training process, thereby improving the model's recognition ability for complex and difficult samples.

[0021] In summary, the method for detecting the posture quality of farmland seedlings by drone based on multimodal perception proposed in the present invention effectively integrates the texture information of RGB images and the spectral information of multispectral images, significantly improving the detection accuracy of seedling targets under complex backgrounds and the recognition ability of seedlings with different postures. The designed MDIPM module and CBAF module can effectively suppress background interference, highlight the characteristics of seedlings, and improve the effectiveness of feature representation; the designed online difficult example mining strategy for seedlings enhances the robustness and generalization ability of the model. This method provides a reliable technical means for realizing the automated and precise evaluation of the posture quality of farmland seedlings, has broad application prospects, can provide timely and accurate field information for agricultural producers, guide agricultural production management, and improve production efficiency and economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the structure diagram of the deep learning model of the present invention, m 1 It is the spectral image of 450nm band, m 2 It is the spectral image of 660nm band, m 3 It is the spectral image of 750nm band, m 4 It is a spectral image of the 840nm band;

[0023] Figure 2 is a flow chart of the method of the present invention;

[0024] Figure 3 This is a structural diagram of the multi-domain interaction prompt module of the present invention. DETAILED DESCRIPTION

[0025] Specific implementation method 1: Combination Figure 1 , Figure 2 The present embodiment is described. The present embodiment is a method for detecting the quality of rice transplanting posture based on multimodal perception, which specifically comprises the following steps:

[0026] The present invention proposes a method for detecting the posture quality of rice seedlings in farmland by using drones based on multimodal perception, aiming to overcome the problems of the prior art, such as insufficient target recognition accuracy in complex farmland backgrounds, low utilization of single modal information, and limited ability to distinguish subtle differences in rice seedlings with different postures. The core of the present invention lies in the effective fusion of multimodal information and the enhancement of rice seedling posture features, thereby achieving high-precision and high-robustness rice seedling posture quality detection. The following are the key points and points to be protected of the present invention:

[0027] 1. Multi-domain interactive prompt module. In response to the problem of complex farmland background interference, the present invention innovatively designs a multi-domain interactive prompt module. This module extracts image information from multiple angles such as frequency domain saliency, color channel characteristics, and learnable weed suppression, and generates a variety of attention masks or weight maps. Through the interactive fusion of multi-domain information, the network is effectively guided to focus on potential seedling areas and suppress the interference of background noise, thereby improving the model's seedling target detection accuracy and recall rate under complex backgrounds.

[0028] 2. Cross-band attention fusion module. In order to fully tap the complementary information of RGB images and multispectral images and improve the ability to distinguish seedlings with different postures, the present invention proposes a cross-band attention fusion module. This module uses specific bands of multispectral images (such as near infrared, red light, etc.) to extract features closely related to different seedling states (such as floating seedlings and injured seedlings), and effectively fuses them with RGB features through a cross-attention mechanism. This module enables the network to pay attention to the subtle differences in the spectral dimensions of seedlings with different postures, significantly enhancing the model's recognition accuracy for abnormal seedlings such as floating seedlings and injured seedlings.

[0029] 3. A seedling posture quality detection framework integrating a multi-domain interactive prompt module and a cross-band attention fusion module. The present invention effectively integrates the proposed multi-domain interactive prompt module and the cross-band attention fusion module into the target detection framework, and constructs an end-to-end multimodal seedling posture quality detection model. The multi-domain interactive prompt module suppresses background interference in the early stage of feature extraction and focuses on potential seedling areas; the cross-band attention fusion module fuses multispectral information in the intermediate feature layer to enhance the ability to distinguish seedlings with different postures. This phased and progressive multimodal information fusion strategy gives full play to the advantages of different modal data and significantly improves the overall detection performance.

[0030] 4. Adaptive model training method based on online difficult example mining strategy. In order to improve the robustness and generalization ability of the model, especially the performance in complex scenarios and difficult samples, the present invention proposes an adaptive model training method based on online difficult example mining strategy. This method dynamically evaluates the difficulty of samples during training and gives higher weights to difficult samples, guiding the model to pay more attention to difficult samples, thereby improving the learning efficiency of the model and the final detection performance.

[0031] In summary, the key point of the present invention is that for the task of detecting the quality of the posture of farmland seedlings, a multi-domain interactive prompt module and a cross-band attention fusion module are innovatively proposed, and they are effectively integrated into the deep learning target detection framework. At the same time, an online difficult case mining strategy is designed to guide model training, further improving the performance of the model. The combination of these key technologies enables the present invention to effectively solve the problem of seedling posture quality detection in complex farmland environments.

[0032] Step S1, collecting RGB images and multispectral images of farmland during the seedling growth period;

[0033] Step S2, preprocessing the RGB image and multispectral image of the farmland during the seedling growth period to obtain the preprocessed RGB image and multispectral image of the farmland during the seedling growth period;

[0034] Step S3, constructing a deep learning model, inputting the pre-processed RGB images and multispectral images of the farmland during the seedling growth period into the deep learning model for training, and obtaining a trained deep learning model;

[0035] Step S4: classify and predict the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested based on the trained deep learning model, and predict the category of the seedlings in the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested as qualified seedlings, injured seedlings or floating seedlings.

[0036] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that: in step S1, RGB images and multispectral images of farmland during the seedling growth period are collected; the specific process is:

[0037] Select drones with high-precision positioning and attitude control to ensure stability during flight and accuracy of data collection;

[0038] Integrate visible light cameras (RGB) and multispectral cameras into UAV platforms;

[0039] The multispectral camera needs to be able to cover the key bands for seedling growth (covering 450nm, 660nm, 750nm, and 840nm bands);

[0040] The key wavebands for seedling growth are 450nm, 660nm, 750nm and 840nm;

[0041] Set the drone's flight altitude and camera parameters;

[0042] Synchronously collect RGB and multispectral images during the seedling growth period

[0043] Plan reasonable flight routes according to the scope and terrain conditions of the farmland to be inspected, ensuring that the image coverage rate is not less than 90% and the lateral overlap rate is not less than 75%;

[0044] Set the drone's flight altitude and camera parameters to obtain image data that meets the resolution requirements (resolution must be at least 2 cm / pixel) to ensure that individual seedlings can be effectively identified;

[0045] The RGB images and multispectral images of the key stages of the seedling growth period (tillering stage or panicle differentiation stage) were collected simultaneously, when the posture characteristics of the seedlings were more obvious;

[0046] Ensure the temporal alignment of data of different modalities and provide a basis for subsequent multimodal information fusion.

[0047] The other steps and parameters are the same as those in the first embodiment.

[0048] Specific implementation method three: This implementation method is different from specific implementation methods one or two in that: in S2, the RGB image and multispectral image of the farmland in the seedling stage are preprocessed to obtain the preprocessed RGB image and multispectral image of the farmland in the seedling stage; the specific process is:

[0049] S21, image geometric correction and atmospheric correction; the specific process is:

[0050] (1) Perform geometric correction on the collected RGB images and multispectral images to eliminate the geometric deformation caused by sensor posture changes and lens distortion, and restore the true geometric shape of the image;

[0051] The geometric correction is implemented based on the camera calibration parameters and the UAV attitude data using the spatial resection method.

[0052] (2) Perform atmospheric correction on the geometrically corrected RGB images and multispectral images to eliminate the influence of atmospheric absorption, scattering and other factors on the spectral information of the ground objects and obtain the real surface reflectance data; it is expressed as:

[0053]

[0054] in,

[0055] L λ is the spectral radiance of the RGB image, in W / (m2 ·sr·nm), W is watt (unit of power), m 2 is square meter (unit of area), sr is steradian (unit of solid angle), nm is nanometer (unit of wavelength), and · is the multiplication sign;

[0056] d is the distance between the sun and the earth, in astronomical units (AU), which is the standard unit for measuring the average distance between the sun and the earth in astronomy;

[0057] E sunλ is the average extraterrestrial irradiance of the sun in a given wavelength band, in W / (m 2 nm);

[0058] θ z is the solar zenith angle;

[0059] ρ λ is the surface reflectivity;

[0060] S22, band registration is performed on the multispectral image after atmospheric correction to obtain a multispectral image after band registration; the specific process is as follows:

[0061] The different bands of the atmospherically corrected multispectral images are registered to ensure that the same object corresponds to the same spatial position in images of different bands, providing a basis for subsequent band information fusion;

[0062] The registration method uses the method of manually selecting control points;

[0063] S23, labeling the RGB image after atmospheric correction and the multispectral image after band registration; the specific process is as follows:

[0064] Label the RGB images after atmospheric correction and the multispectral images after band registration as qualified seedlings, injured seedlings or floating seedlings;

[0065] Use the interactive labeling tool X-AnyLabeling to label the quality status information of each seedling in the farmland (injured seedlings, floating seedlings, qualified seedlings).

[0066] S24, image enhancement and denoising processing; the specific process is:

[0067] (1) To address the problems of low contrast and blur in farmland images, the CLAHE (Contrast Limited Adaptive Histogram Equalization) image enhancement algorithm is used to enhance the annotated RGB images and multispectral images to obtain enhanced images and improve the visibility of seedling targets.

[0068] (2) De-noising the enhanced image to obtain a denoised image. The specific process is as follows:

[0069] The enhanced image is convolved with the Gaussian kernel to obtain the denoised image.

[0070] The Gaussian filter denoising algorithm is used to remove noise from the image and improve the robustness of feature extraction. The denoising algorithm used is expressed by the following formula:

[0071]

[0072] Among them, σ represents the standard deviation of the Gaussian kernel, (x, y) represents the coordinate position of the pixel in the image, and G(x, y) represents the Gaussian kernel;

[0073] The denoised image I'(x,y) is the convolution of the original image I(x,y) and the Gaussian kernel G(x,y):

[0074]

[0075] Among them, k represents the radius of the Gaussian kernel, (i, j) represents the relative coordinate index inside the Gaussian kernel, and G(i, j) represents the Gaussian kernel function.

[0076] The other steps and parameters are the same as those in the first or second embodiment.

[0077] Specific implementation method 4: This implementation method is different from any one of the specific implementation methods 1 to 3 in that: a deep learning model is constructed in S3; the specific process is:

[0078] The deep learning model consists of the Backbone part, the Neck part, and the prediction part Head;

[0079] The Backbone part includes the first convolution layer, the second convolution layer, the first C3k2 module, the third convolution layer, the second C3k2 module, the fourth convolution layer, the third C3k2 module, the fifth convolution layer, the fourth C3k2 module, the SPPF module, and the C2PSA module in sequence;

[0080] The Neck part includes the first upsampling UpSample, the first fully connected layer Concat, the first cross-band attention fusion module CBAF, the first multi-domain interactive prompt module MDIPM, the fifth C3k2 module, the second upsampling UpSample, the second fully connected layer Concat, the sixth C3k2 module, the third upsampling UpSample, the third fully connected layer Concat, the seventh C3k2 module, the sixth convolutional layer, the fourth fully connected layer Concat, the eighth C3k2 module, the seventh convolutional layer, the fifth fully connected layer Concat, the ninth C3k2 module, the eighth convolutional layer, the sixth fully connected layer Concat, and the tenth C3k2 module;

[0081] The prediction part Head includes a first prediction head, a second prediction head, a third prediction head, and a fourth prediction head.

[0082] The other steps and parameters are the same as those in Specific Embodiments 1 to 3.

[0083] Specific implementation method 5: This implementation method is different from the specific implementation methods 1 to 4 in that the specific working process of the Backbone part is:

[0084] The RGB image of the farmland in the seedling stage preprocessed in step S2 is sequentially input into the first convolution layer, the second convolution layer, and the first C3k2 module, and the first C3k2 module outputs a feature map A;

[0085] Feature map A is sequentially input into the third convolutional layer and the second C3k2 module, and the second C3k2 module outputs feature map B;

[0086] Feature map B is sequentially input into the fourth convolutional layer and the third C3k2 module, and the third C3k2 module outputs feature map C;

[0087] The feature map C is sequentially input into the fifth convolutional layer and the fourth C3k2 module, and the fourth C3k2 module outputs the feature map D;

[0088] Feature map D is input into the SPPF module and the C2PSA module in sequence, and the C2PSA module outputs feature map E.

[0089] The other steps and parameters are the same as those in Specific Embodiments 1 to 4.

[0090] Specific implementation method 6: This implementation method is different from the specific implementation methods 1 to 5 in that the specific working process of the Neck part is:

[0091] The feature map E undergoes the first upsampling UpSample, and the first upsampling UpSample outputs the feature map F;

[0092] Feature map C and feature map F are concatted through the first fully connected layer, and the first fully connected layer concat outputs feature map G;

[0093] The feature map B passes through the first cross-band attention fusion module CBAF, and the first cross-band attention fusion module CBAF outputs the feature map H;

[0094] The feature map A and the RGB image of the farmland in the seedling stage preprocessed in step S2 are passed through a first multi-domain interactive prompt module MDIPM, and the first multi-domain interactive prompt module MDIPM outputs a feature map I;

[0095] The feature map G passes through the fifth C3k2 module, and the fifth C3k2 module outputs the feature map J;

[0096] The feature map J undergoes a second upsampling UpSample, and the second upsampling UpSample outputs a feature map J′;

[0097] The feature map J′, the feature map B, and the feature map H are concatted through the second fully connected layer, and the second fully connected layer concat outputs the feature map K;

[0098] The feature map K passes through the sixth C3k2 module, and the sixth C3k2 module outputs the feature map L;

[0099] The feature map L undergoes the third upsampling UpSample, and the third upsampling UpSample outputs the feature map M;

[0100] Step S2: The RGB image of the farmland in the seedling stage, the feature map M, the feature map A, and the feature map I after preprocessing are concatted by the third fully connected layer, and the third fully connected layer concat outputs the feature map N;

[0101] The feature map N is input into the seventh C3k2 module, and the seventh C3k2 module outputs the feature map O;

[0102] The feature map O is input into the sixth convolutional layer, and the sixth convolutional layer outputs the feature map O′;

[0103] The feature map O′ and the feature map L are sequentially input into the fourth fully connected layer Concat and the eighth C3k2 module, and the eighth C3k2 module outputs the feature map P;

[0104] The feature map P is input into the seventh convolutional layer, and the seventh convolutional layer outputs the feature map Q;

[0105] The seventh convolutional layer output feature map Q and the fifth C3k2 module output feature map J are sequentially input into the fifth fully connected layer Concat and the ninth C3k2 module, and the ninth C3k2 module outputs feature map R;

[0106] The feature map R is input into the eighth convolutional layer, and the eighth convolutional layer outputs the feature map S;

[0107] The feature map S and the feature map E are input into the sixth fully connected layer Concat, the sixth fully connected layer Concat outputs the feature map and inputs the tenth C3k2 module, and the tenth C3k2 module outputs the feature map T.

[0108] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 5.

[0109] Specific implementation method 7: This implementation method is different from the specific implementation methods 1 to 6 in that the specific working process of the prediction part Head is:

[0110] The feature map T output by the tenth C3k2 module is input into the fourth prediction head, and the fourth prediction head outputs the classification result;

[0111] The feature map R output by the ninth C3k2 module is input into the third prediction head, and the third prediction head outputs the classification result;

[0112] The feature map P output by the eighth C3k2 module is input into the second prediction head, and the second prediction head outputs the classification result;

[0113] The feature map O output by the seventh C3k2 module is input into the first prediction head, and the first prediction head outputs the classification result;

[0114] Each of the first prediction head, the second prediction head, the third prediction head, and the fourth prediction head outputs a prediction result including a bounding box position, a confidence level, and a category probability.

[0115] The multi-scale feature map can obtain the detection results of the seedling posture quality (normal, floating seedlings, and damaged seedlings) through the prediction head.

[0116] Each prediction head is responsible for detecting seedling targets of a specific scale. Feature maps of larger scales are used to detect larger seedling targets, while feature maps of smaller scales are used to detect smaller seedling targets.

[0117] like Figure 1 As shown, the prediction part contains prediction heads (Detect) of multiple scales.

[0118] Input the preprocessed RGB image into the Backbone part of the YOLOv11 object detection framework for feature extraction;

[0119] The backbone part usually consists of a series of convolutional layers, pooling layers, and activation functions. For example, the C3K2 module is used to extract shallow texture features and deep semantic features of images.

[0120] Multi-scale feature maps are extracted at different stages of Backbone for subsequent multi-scale feature fusion and object detection.

[0121] The other steps and parameters are the same as those in Specific Embodiments 1 to 6.

[0122] Specific implementation example eight: This implementation example is different from any one of specific implementation examples one to seven in that: the specific working process of the first multi-domain interaction prompt module MDIPM is as follows:

[0123] In order to solve the problem of complex background interference in seedling target detection, this paper designs a novel multi-domain interactive prompt module. The core idea of ​​this module is to extract the saliency information and attention weight of the image from different angles, guide the network to pay more attention to the potential seedling area, and suppress the interference of background noise. The structure of this module is as follows Figure 3 shown.

[0124] Specifically, the multi-domain interactive prompt module contains the following three sub-modules:

[0125] 1) Obtain the saliency mask attention map; the specific process is: this submodule aims to extract the saliency information of the image from the frequency domain perspective.

[0126] 11) The image As input;

[0127] image Output feature map A for the first C3k2 module and the RGB image of the farmland in the seedling stage after preprocessing in step S2;

[0128] Among them, R represents a real number, H represents the height of the image, and W represents the width of the image;

[0129] 12) The image Convert to grayscale I gray ∈R H×W ;

[0130] image Output feature map A for the first C3k2 module and the RGB image of the farmland in the seedling stage after preprocessing in step S2;

[0131] 13) For grayscale image I gray Perform a two-dimensional discrete Fourier transform to obtain the grayscale image F(u,v) after the two-dimensional discrete Fourier transform:

[0132] F(u,v)=F(I gray )(u,v)

[0133] Among them, F(I gray ) represents the grayscale image I gray Perform a two-dimensional discrete Fourier transform operation; (u, v) represents the grayscale image I gray The frequency domain coordinates of; F(u,v) represents the grayscale image I gray Grayscale image after two-dimensional discrete Fourier transform;

[0134] 14) Perform a center shift on the grayscale image F(u,v) after the two-dimensional discrete Fourier transform to obtain the spectrum F shift (u,v);

[0135] 15) Calculate the amplitude spectrum of the spectrum M(u,v) = 20×log(|F shift (u,v)|+1), for visualization;

[0136] In order to highlight the high-frequency information, an ideal high-pass filter mask Mask(u,v) with a zero center area is constructed:

[0137]

[0138] Where r represents the radius of the circular area in the center of the mask;

[0139] 16) Apply the mask Mask(u,v) to the spectrum F shift (u,v) gets the spectrum F after applying the mask masked (u,v)=F shift (u,v)×Mask(u,v);

[0140] 17) Restore the saliency map in the spatial domain through inverse Fourier transform

[0141] in, represents the inverse Fourier transform, I saliency (x,y) represents the saliency map;

[0142] 18) For the saliency map I saliency (x,y) is normalized to get I saliency_norm (x,y);

[0143] 19) According to the threshold τ saliency Generate a binary saliency mask attention map M saliency :

[0144]

[0145] Among them, M saliency (x, y) represents the binary saliency mask attention map, M saliency ∈R H×W×1 ;

[0146] Binarized saliency mask M saliency (x, y) can effectively highlight the high-frequency information in the image, corresponding to the area with drastic changes in the image, such as the edge of the seedling.

[0147] 2) Obtain the color attention weight map; the specific process is:

[0148] Considering that seedlings usually present a specific green tone, this submodule aims to use color information to guide the network to focus on potential seedling areas.

[0149] 21) Convert the RGB image of the rice seedling stage farmland preprocessed in step S2 into the HSV color space to obtain the image I in the HSV color space. HSV ∈R H×W×3 ;

[0150] 22) Define the learnable hue range θ H =[Hmin ,H max ], saturation range θ S =[S min ,S max ], and the brightness range θ V =[V min ,V max ];

[0151] Among them, H min Indicates the minimum hue value; H max Indicates the maximum value of hue; S min Indicates the minimum saturation value; S max Indicates the maximum saturation value; V min Indicates the minimum brightness; V max Indicates the maximum brightness;

[0152] These thresholds are learned and optimized dynamically by the model during training;

[0153] 23) Calculate the image I in HSV color space HSV The probability mask P of each pixel belonging to the seedling area color :

[0154] P color (x,y)=σ(w H ×(I HSV (x,y,0)-H center ) 2 +w S ×(I HSV (x,y,1)-S center ) 2 +w V ×(I HSV (x,y,2)-V center ) 2 )

[0155] in,

[0156] H center is the hue range θ H The central value of S center is the saturation range θ S The center value of V center They are the brightness range θ V The central value of

[0157] I HSV (x,y,0) represents image I HSV The hue component of the pixel (x, y) in the HSV color space;

[0158] I HSV (x,y,1) represents image IHSV The saturation component of the HSV color space of the pixel (x, y);

[0159] I HSV (x,y,2) represents image I HSV The brightness component of the HSV color space of the pixel (x, y);

[0160] σ(·) is the Sigmoid function;

[0161] w H ,w S ,w V is the learnable weight;

[0162] 24) For the probability mask P color Normalize and get the color attention weight map A color :

[0163]

[0164] Among them, min(P color ) is the probability mask P color The minimum value of max(P color ) is the probability mask P color The maximum value of

[0165] 25) In order to increase smoothness, the color attention weight map A color Perform Gaussian filtering to obtain the color attention weight map A after Gaussian filtering. color ;

[0166] 3) Obtain the attention-guiding feature map; the specific process is:

[0167] In order to further suppress potential interference factors such as weeds, this submodule introduces a learnable attention branch.

[0168] 31) The RGB image of the rice seedling stage farmland preprocessed in step S2 is passed through a convolution block Conv weed Generate attention weight map A weed ; expressed as:

[0169] A weed =σ(Conv weed (F shallow )),

[0170] Among them, σ is the Sigmoid activation function;

[0171] Conv weed Convolution block, Conv weed It includes convolution layer, maximum pooling layer, and convolution layer in sequence; Aweed ∈R H ×W×1 ;

[0172] F shallow The RGB image of the farmland at the seedling stage after preprocessing in step S2;

[0173] Shallow features contain richer texture and edge information, which helps to distinguish weeds from seedlings. 32),

[0175] The saliency mask attention map M saliency After passing through two convolutional layers in sequence, the feature map F is obtained. saliency ;

[0176] The color attention weight map A after Gaussian filtering color After passing through two convolutional layers in sequence, the feature map F is obtained. color ;

[0177] The attention weight map A weed After passing through two convolutional layers in sequence, the feature map F is obtained. weed ;

[0178] 33) The feature map F saliency , feature map F color , feature map F weed Splicing in the channel dimension, and then passing through the linear layer to obtain the feature map after dimensionality reduction;

[0179] The feature map after dimensionality reduction and the feature map A output by the first C3k2 module are input into the cross attention module for fusion to obtain the final attention-guided feature map

[0180] The other steps and parameters are the same as those in Specific Embodiments 1 to 7.

[0181] Specific implementation method 9: This implementation method is different from any one of specific implementation methods 1 to 8 in that: the specific working process of the first cross-band attention fusion module CBAF is:

[0182] The designed cross-band attention fusion module (CBAF) aims to fuse RGB image features with the specific band information provided by multispectral images to enhance the recognition ability of seedlings with different postures.

[0183] The input of the module includes: the feature map output by the second C3k2 module And the pre-processed 450nm, 660nm, 750nm, 840nm multi-band spectral images; C 3 Indicates the number of channels of feature map B, W3 represents the width of feature map B, H 3 Represents the height of feature map B;

[0184] The specific steps are as follows: 1)

[0186] The multispectral images of 450nm and 660nm bands are spliced ​​in the channel dimension and then input into the seedling state feature extractor 1 (SCFE1) for feature extraction to obtain the damaged seedling feature D 1 ,

[0187] The main function of SCFE1 is to fuse, extract features and reduce the dimension of dual-band information so that its feature dimension is consistent with the feature map B output by the second C3k2 module.

[0188] Seedling condition feature extractor 1 is SCFE1 (Seedling Condition Feature Extractor 1);

[0189] SCFE1 includes the first convolution block, the second convolution block, and the third convolution block in sequence;

[0190] Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a ReLU activation function layer, a convolution layer, a ReLU activation function layer, a convolution layer, and a maximum pooling layer in sequence;

[0191] The blue light (450nm) and red light (660nm) bands are more sensitive to chlorophyll absorption and can effectively reflect chlorophyll loss, so they are more suitable for extracting damaged seedling characteristics. 2)

[0193] The multispectral images of 750nm and 840nm bands are spliced ​​in the channel dimension and input into the seedling state feature extractor 2 (with the same structure as the seedling state feature extractor 1) for feature extraction to obtain the floating seedling feature D 2 ,

[0194]

[0195] Seedling condition feature extractor 2 is SCFE2 (Seedling Condition Feature Extractor 2);

[0196] SCFE2 includes the first convolution block, the second convolution block, and the third convolution block in sequence;

[0197] Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a ReLU activation function layer, a convolution layer, a ReLU activation function layer, a convolution layer, and a maximum pooling layer in sequence;

[0198] SCFE2 (Seedling Condition Feature Extractor 2, SCFE2) is designed to perform dual-band fusion, feature extraction and dimensionality reduction. The near-infrared band (750nm and 840nm) is sensitive to changes in the cell structure and moisture content of vegetation and is more suitable for extracting floating seedling features.

[0199] 3) Cross-Attention:

[0200] The characteristics of the injured seedlings D 1 As the key K (Key), the feature map B output by the second C3k2 module is used as the query Q (Query) and the value V (Value) respectively, and the cross-attention fusion operation is performed to obtain the feature map

[0201] Such an operation enables the RGB features to focus on the spectral information related to damaged seedlings, thereby enhancing the network's ability to identify damaged seedlings.

[0202] The floating feature D 2 As the key K (Key), the feature map B output by the second C3k2 module is used as the query Q (Query) and the value V (Value), and the cross-attention fusion operation is performed to obtain the feature map

[0203] This operation enables the RGB features to focus on the spectral information related to floating rice seedlings, thus improving the network's ability to identify floating rice seedlings.

[0204] The calculation formula of the cross-attention fusion operation is as follows:

[0205]

[0206] Among them, d k is the dimension of the key K(Key); the superscript T indicates transposition; softmax indicates the activation function; Attention(Q,K,V) indicates cross-attention fusion operation;

[0207] 4) The feature map obtained by cross-attention fusion operation and Fusion is performed to obtain the cross-band attention map F CBAF ; The specific process is:

[0208] Fusion is channel concatenation followed by dimensionality reduction through a convolutional layer:

[0209]

[0210] Among them, Conv fuse is the convolutional layer used to process the fused feature map, Concat() is concatenation, and F CBAF is the cross-band attention map;

[0211] 5) The cross-band attention map F CBAF The Hadamard product is performed with the feature map B output by the second C3k2 module to obtain the features after attention

[0212]

[0213] Among them, ⊙ is the Hadamard product.

[0214] like Figure 1 As shown in the figure, in the fusion layer, the feature pyramid network (FPN) structure is used to perform multi-scale feature fusion;

[0215] The feature maps from different stages of Backbone are fused through upsampling and concatenation operations to obtain feature maps containing information of different scales.

[0216] The specific fusion method includes two pathways: bottom-up and top-down, and features of the same scale are fused through lateral connections.

[0217] The other steps and parameters are the same as those in Specific Implementation 1 to 8-1.

[0218] Specific implementation method ten: This implementation method is different from any one of specific implementation methods one to nine in that: the deep learning model is trained to obtain a trained deep learning model; the specific process is:

[0219] Using the designed online hard example mining strategy to guide model adaptive training includes the following aspects:

[0220] (1) Input the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 into the deep learning model, and the first prediction head of the deep learning model outputs the category prediction probability distribution p;

[0221] Calculate the cross entropy loss for each sample based on the true label y and the predicted probability distribution p output by the first prediction head The calculation formula is as follows:

[0222]

[0223] Among them, C is the number of categories, and its value is 3;

[0224] y i is the true label of each sample belonging to the i-th category;

[0225] p i is the probability p of each sample belonging to the i-th category predicted by the deep learning model i ;

[0226] i=1, 2 or 3, i represents qualified seedlings, injured seedlings or floating seedlings; (2)

[0228] In step S2, the RGB image and multispectral image of the farmland in the seedling stage after preprocessing are input into the deep learning model. The first prediction head of the deep learning model outputs the category prediction probability distribution p, and calculates the prediction entropy H(p) of each sample to measure the uncertainty of the model prediction. The higher the prediction entropy, the more uncertain the model's prediction of the sample. The calculation formula is as follows:

[0229]

[0230] In addition, in order to evaluate the stability of model prediction, the Dropout (random inactivation) technique is used to perform T forward propagations to obtain T prediction probability distributions p (1) ,p (2) ,...,p (T) , calculate the prediction variance Var(p) of each sample, the calculation formula is as follows:

[0231]

[0232] Among them, Var(·) represents variance calculation;

[0233] It represents the probability that each sample predicted by the first prediction head of the deep learning model belongs to the i-th category when the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the first time;

[0234] It indicates that the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the second time, and the probability that each sample belongs to the i-th category predicted by the first prediction head of the deep learning model;

[0235] represents the probability that each sample predicted by the first prediction head of the deep learning model belongs to the i-th category when the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the Tth time;

[0236] (3) Construct a scoring function for each sample:

[0237] Taking into account the cross entropy loss and uncertainty measurement, a scoring function S is designed to evaluate the difficulty of each sample; the scoring function is as follows:

[0238]

[0239] Among them, α, β, and γ are hyperparameters used to balance the weights of the loss function and different uncertainty metrics, and they are set to 1.0, 0.1, and 0.1 respectively;

[0240] (4) Calculate the score S of all samples and select the K% of samples with the highest scores as hard examples based on the scores;

[0241] K is a preset fixed value, which is empirically set to 30% in this method.

[0242] (5) For samples selected as difficult examples, a higher weight is given when calculating the loss function, so that the model pays more attention to the learning of these difficult examples;

[0243] The weight function w(S) associated with the score of hard examples is defined as follows:

[0244] w(S)=1+λ×tanh(S)

[0245] Among them, λ is a hyperparameter that controls the magnitude of weight adjustment;

[0246] For each hard example, the gradient is calculated using the weighted loss function:

[0247]

[0248] in,

[0249] S(p i ) is the score of the difficult sample calculated using the scoring function;

[0250] w(S(p i )) is the weight of the hard sample;

[0251] For each non-hard example, the gradient is calculated using a weighted loss function with a weight of 1:

[0252]

[0253] Use the back-propagation algorithm to update deep learning model parameters;

[0254] (6) Train the deep learning model until the weighted loss function Converge and obtain a trained deep learning model.

[0255] The other steps and parameters are the same as those in Specific Implementation Methods 1 to 9.

[0256] Although the technical solution of the present invention can effectively achieve the established goals, there may be some alternative solutions when the performance requirements are adjusted. However, these alternative solutions are often difficult to achieve the technical effects of the present invention in dealing with complex farmland environments and making full use of multimodal information.

[0257] The core technical challenge of the present invention is how to effectively fuse RGB and multispectral information and accurately identify seedlings of different postures in a complex background. The present invention effectively solves this series of problems by constructing a multi-domain interactive prompt module and a cross-band attention fusion module, and combining it with the designed online difficult example mining strategy. The multi-domain interactive prompt module aims to suppress background interference and highlight the seedling target; the cross-band attention fusion module aims to fuse multispectral information and enhance the ability to distinguish different postures; the online difficult example mining strategy improves the robustness and generalization of the model.

[0258] If an alternative to S4 is considered, the most direct idea is to rely only on data from a single modality for seedling posture quality detection. For example, only RGB images collected by drones are used for detection. Under this scheme, traditional image processing-based methods, such as color segmentation, texture analysis, etc., can be used in combination with machine learning classifiers to identify and classify seedling targets. However, relying solely on RGB images for detection is easily disturbed by factors such as lighting changes, shadows, and weeds, making it difficult to accurately distinguish between seedlings of different postures, especially some injured or slightly floating seedlings whose apparent differences are not obvious. In contrast, the present invention integrates multispectral information and utilizes the differences in spectral reflectance of seedlings of different postures to more effectively improve recognition accuracy.

[0259] In terms of multimodal information fusion, although there are some other fusion strategies, such as simple feature splicing or weighted fusion, these methods may not be able to fully explore the deep correlation and complementary information between different modal data. Feature splicing simply connects the feature vectors of different modalities together, lacking interaction between modalities; weighted fusion requires manual setting of weights, and it is difficult to adaptively adjust the importance of different modalities. In contrast, the cross-band attention fusion module (CBAF) designed in the present invention dynamically learns the contribution of different spectral bands to seedling posture discrimination through the attention mechanism, and effectively fuses it with the RGB features, which can make more refined use of multimodal information.

[0260] In addition, in the seedling posture quality detection link of step S5, although there are other difficult example mining strategies, such as the fixed threshold difficult example selection method based on cross entropy loss, this method lacks consideration of the uncertainty of model prediction in the way of evaluating the difficulty of samples. In contrast, the online difficult example mining strategy adopted by the present invention can dynamically select difficult examples according to the uncertainty of the model's prediction of samples, and give them higher training weights, so that the model can pay more attention to samples that are difficult to learn, thereby more effectively improving the robustness and generalization ability of the model.

[0261] In summary, although there are alternative solutions in some aspects of the technical solution, these alternative solutions may have certain limitations in solving the interference of complex farmland background, effectively integrating multimodal information, and improving the model's fine-grained recognition ability of seedlings with different postures, and it is difficult to achieve the technical effect of the present invention. The multimodal perception method proposed in the present invention, as well as the specific module design and training strategy, are proposed to more effectively solve the specific problem of farmland seedling posture quality detection, and are necessary and innovative.

[0262] What are the advantages of this invention compared with the closest prior art?

[0263] At present, the technology for detecting the posture quality of farmland seedlings mainly focuses on manual inspection and analysis methods based on single-modal remote sensing images. Compared with these existing technologies, the method for detecting the posture quality of farmland seedlings by drone based on multi-modal perception proposed in the present invention has significant advantages in terms of the depth of information fusion, detection accuracy and robustness, and intelligent level, and overcomes many limitations of the existing technologies, which are specifically reflected in the following aspects:

[0264] 1. Significant improvement in the accuracy of seedling target detection under complex farmland background: Existing seedling detection methods, especially those that rely only on RGB images, often show low detection accuracy when faced with complex and changeable farmland environments. Factors such as soil, water, weeds, and light changes can easily cause serious interference, leading to false detection and missed detection of seedling targets. Traditional image processing technology is difficult to effectively distinguish these interference factors from seedling targets, and general target detection models directly applied to RGB images are also difficult to fully utilize prior knowledge of farmland scenes. In contrast, the present invention innovatively proposes a multi-domain interactive prompt module. This module extracts image saliency information from multiple dimensions such as frequency domain and color space, and incorporates a learnable weed suppression mechanism to effectively guide the network to focus on potential seedling areas and significantly suppress the interference of background noise.

[0265] 2. Fine seedling posture quality recognition capability: Traditional seedling posture quality detection methods based on RGB images face challenges in distinguishing seedlings of different postures, especially for some cases with subtle apparent differences, such as slightly injured seedlings and healthy seedlings, which are difficult to accurately distinguish based on RGB information alone. The present invention, by fusing multispectral information, can utilize the reflectivity differences of seedlings with different postures in different spectral bands to achieve more refined posture quality recognition. For example, there are significant differences in the reflectivity of healthy seedlings and damaged seedlings in the near-infrared band, and the spectral reflectance characteristics of floating seedlings are different from those of healthy seedlings due to changes in water status. The cross-band attention fusion module proposed in the present invention can effectively capture these subtle spectral differences, combined with the texture and shape information provided by RGB images, so as to more accurately distinguish floating seedlings, injured seedlings and reasonable seedlings, and improve the accuracy and reliability of posture quality assessment, which is difficult to achieve with existing technologies that rely solely on RGB images.

[0266] 3. Improvement of model robustness and generalization ability: Traditional deep learning model training usually adopts uniform sampling or random sampling, which easily causes the model to spend too much energy on a large number of simple samples, but insufficient learning of difficult examples that are not highly distinguishable and easily confused, resulting in limited generalization ability of the model in complex scenarios. To this end, the present invention designs a strategy for online difficult example mining of rice seedlings. During the training process, this strategy dynamically identifies "difficult example" samples that are difficult for the model to accurately identify (such as samples of some floating seedlings and damaged seedlings that are difficult to distinguish), and assigns these samples higher training weights, forcing the model to pay more attention to the learning of these difficult examples. This adaptive training method effectively improves the robustness and generalization ability of the model and improves the practical application value of the model.

[0267] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention, but these corresponding changes and modifications should all fall within the scope of protection of the claims attached to the present invention.

Claims

1. A rice transplanting posture quality detection method based on multimodal perception, characterized in that: The specific process of the method is: Step S1, collecting RGB images and multispectral images of farmland during the seedling growth period; Step S2, preprocessing the RGB image and multispectral image of the farmland during the seedling growth period to obtain the preprocessed RGB image and multispectral image of the farmland during the seedling growth period; Step S3, constructing a deep learning model, inputting the pre-processed RGB images and multispectral images of the farmland during the seedling growth period into the deep learning model for training, and obtaining a trained deep learning model; Step S4: classify and predict the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested based on the trained deep learning model, and predict the category of the seedlings in the RGB images and multispectral images of the farmland during the growth period of the seedlings to be tested as qualified seedlings, injured seedlings or floating seedlings.

2. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 1, characterized in that: In step S1, RGB images and multispectral images of farmland during the rice seedling growth period are collected; the specific process is as follows: Integrate visible light cameras and multispectral cameras into UAV platforms; RGB images and multispectral images of the seedling growth period are collected simultaneously.

3. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 2, characterized in that: In the S2, the RGB image and the multispectral image of the farmland in the seedling stage are preprocessed to obtain the preprocessed RGB image and the multispectral image of the farmland in the seedling stage; the specific process is: S21, image geometric correction and atmospheric correction; the specific process is: (1) Perform geometric correction on the collected RGB images and multispectral images; (2) Perform atmospheric correction on the geometrically corrected RGB image and multispectral image to obtain the real surface reflectance data; expressed as: in, Lλ is the spectral radiance of the RGB image; d is the distance between the sun and the earth; E sunλ is the average extraterrestrial irradiance of the sun in a given band; θ z is the solar zenith angle; ρ λ is the surface reflectivity; S22, band registration is performed on the multispectral image after atmospheric correction to obtain a multispectral image after band registration; the specific process is as follows: The different bands of the atmospherically corrected multispectral image are registered to ensure that the same object corresponds to the same spatial position in images of different bands; S23, labeling the RGB image after atmospheric correction and the multispectral image after band registration; the specific process is as follows: Label the RGB images after atmospheric correction and the multispectral images after band registration as qualified seedlings, injured seedlings or floating seedlings; S24, image enhancement and denoising processing; the specific process is: (1) Enhance the annotated RGB image and multispectral image to obtain an enhanced image; (2) De-noising the enhanced image to obtain a denoised image. The specific process is as follows: The enhanced image is convolved with the Gaussian kernel to obtain the denoised image.

4. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 3, characterized in that: The deep learning model is constructed in S3; the specific process is: The deep learning model consists of the Backbone part, the Neck part, and the prediction part Head; The Backbone part includes the first convolution layer, the second convolution layer, the first C3k2 module, the third convolution layer, the second C3k2 module, the fourth convolution layer, the third C3k2 module, the fifth convolution layer, the fourth C3k2 module, the SPPF module, and the C2PSA module in sequence; The Neck part includes the first upsampling UpSample, the first fully connected layer Concat, the first cross-band attention fusion module CBAF, the first multi-domain interactive prompt module MDIPM, the fifth C3k2 module, the second upsampling UpSample, the second fully connected layer Concat, the sixth C3k2 module, the third upsampling UpSample, the third fully connected layer Concat, the seventh C3k2 module, the sixth convolutional layer, the fourth fully connected layer Concat, the eighth C3k2 module, the seventh convolutional layer, the fifth fully connected layer Concat, the ninth C3k2 module, the eighth convolutional layer, the sixth fully connected layer Concat, and the tenth C3k2 module; The prediction part Head includes a first prediction head, a second prediction head, a third prediction head, and a fourth prediction head.

5. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 4, characterized in that: The specific working process of the Backbone part is: The RGB image of the farmland in the seedling stage preprocessed in step S2 is sequentially input into the first convolution layer, the second convolution layer, and the first C3k2 module, and the first C3k2 module outputs a feature map A; Feature map A is sequentially input into the third convolutional layer and the second C3k2 module, and the second C3k2 module outputs feature map B; Feature map B is sequentially input into the fourth convolutional layer and the third C3k2 module, and the third C3k2 module outputs feature map C; The feature map C is sequentially input into the fifth convolutional layer and the fourth C3k2 module, and the fourth C3k2 module outputs the feature map D; Feature map D is input into the SPPF module and the C2PSA module in sequence, and the C2PSA module outputs feature map E.

6. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 5, characterized in that: The specific working process of the Neck part is: The feature map E undergoes the first upsampling UpSample, and the first upsampling UpSample outputs the feature map F; Feature map C and feature map F are concatted through the first fully connected layer, and the first fully connected layer concat outputs feature map G; The feature map B passes through the first cross-band attention fusion module CBAF, and the first cross-band attention fusion module CBAF outputs the feature map H; The feature map A and the RGB image of the farmland in the seedling stage preprocessed in step S2 are passed through a first multi-domain interactive prompt module MDIPM, and the first multi-domain interactive prompt module MDIPM outputs a feature map I; The feature map G passes through the fifth C3k2 module, and the fifth C3k2 module outputs the feature map J; The feature map J undergoes a second upsampling UpSample, and the second upsampling UpSample outputs a feature map J′; The feature map J′, the feature map B, and the feature map H are concatted through the second fully connected layer, and the second fully connected layer concat outputs the feature map K; The feature map K passes through the sixth C3k2 module, and the sixth C3k2 module outputs the feature map L; The feature map L undergoes the third upsampling UpSample, and the third upsampling UpSample outputs the feature map M; Step S2: The RGB image of the farmland in the seedling stage, the feature map M, the feature map A, and the feature map I after preprocessing are concatted by the third fully connected layer, and the third fully connected layer concat outputs the feature map N; The feature map N is input into the seventh C3k2 module, and the seventh C3k2 module outputs the feature map O; The feature map O is input into the sixth convolutional layer, and the sixth convolutional layer outputs the feature map O′; The feature map O′ and the feature map L are sequentially input into the fourth fully connected layer Concat and the eighth C3k2 module, and the eighth C3k2 module outputs the feature map P; The feature map P is input into the seventh convolutional layer, and the seventh convolutional layer outputs the feature map Q; The seventh convolutional layer output feature map Q and the fifth C3k2 module output feature map J are sequentially input into the fifth fully connected layer Concat and the ninth C3k2 module, and the ninth C3k2 module outputs feature map R; The feature map R is input into the eighth convolutional layer, and the eighth convolutional layer outputs the feature map S; The feature map S and the feature map E are input into the sixth fully connected layer Concat, the sixth fully connected layer Concat outputs the feature map and inputs the tenth C3k2 module, and the tenth C3k2 module outputs the feature map T.

7. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 6, characterized in that: The specific working process of the prediction part Head is as follows: The feature map T output by the tenth C3k2 module is input into the fourth prediction head, and the fourth prediction head outputs the classification result; The feature map R output by the ninth C3k2 module is input into the third prediction head, and the third prediction head outputs the classification result; The feature map P output by the eighth C3k2 module is input into the second prediction head, and the second prediction head outputs the classification result; The feature map O output by the seventh C3k2 module is input into the first prediction head, and the first prediction head outputs the classification result; Each of the first prediction head, the second prediction head, the third prediction head, and the fourth prediction head outputs a prediction result including a bounding box position, a confidence level, and a category probability.

8. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 7, characterized in that: The specific working process of the first multi-domain interaction prompt module MDIPM is as follows: 1) Obtain the saliency mask attention map; the specific process is: 11) The image As input; image Output feature map A for the first C3k2 module and the RGB image of the farmland in the seedling stage after preprocessing in step S2; Among them, R represents a real number, H represents the height of the image, and W represents the width of the image; 12) Image Convert to grayscale I gray ∈R H×W ; image Output feature map A for the first C3k2 module and the RGB image of the farmland in the seedling stage after preprocessing in step S2; 13) For grayscale image I gray Perform a two-dimensional discrete Fourier transform to obtain the grayscale image F(u,v) after the two-dimensional discrete Fourier transform: F(u,v)=F(I gray )(u,v) Among them, F(I gray ) represents the grayscale image I gray Perform a two-dimensional discrete Fourier transform operation; (u, v) represents the grayscale image I gray The frequency domain coordinates of; F(u,v) represents the grayscale image I gray Grayscale image after two-dimensional discrete Fourier transform; 14) Perform a center shift on the grayscale image F(u,v) after the two-dimensional discrete Fourier transform to obtain the spectrum F shift (u,v); 15) Construct an ideal high-pass filter mask Mask(u,v) with a zero center area: Where r represents the radius of the circular area in the center of the mask; 16) Apply the mask Mask(u,v) to the spectrum F shift (u,v) gets the spectrum F after applying the mask masked (u,v)=F shift (u,v)×Mask(u,v); 17) Restore the saliency map in the spatial domain through inverse Fourier transform in, represents the inverse Fourier transform, I saliency (x,y) represents the saliency map; 18) For the saliency map I saliency (x,y) is normalized to get I saliency_norm (x,y); 19) According to the threshold τ saliency Generate a binary saliency mask attention map M saliency : Among them, M saliency (x, y) represents the binary saliency mask attention map, M saliency ∈R H×W×1 ; 2) Obtain the color attention weight map; the specific process is: 21) Convert the RGB image of the rice seedling stage farmland preprocessed in step S2 into the HSV color space to obtain the image I in the HSV color space. HSV ∈R H×W×3 ; 22) Define the learnable hue range θ H =[H min ,H max ], saturation range θ S =[S min ,S max ], and the brightness range θ V =[V min ,V max ]; Among them, H min Indicates the minimum hue value; H max Indicates the maximum value of hue; S min Indicates the minimum saturation value; S max Indicates the maximum saturation value; V min Indicates the minimum brightness; V max Indicates the maximum brightness; 23) Calculate the image I in HSV color space HSV The probability mask P of each pixel belonging to the seedling area color : P color (x,y)=σ(w H ×(I HSV (x,y,0)-H center ) 2 +w S ×(I HSV (x,y,1)-S center ) 2 +w V ×(I HSV (x,y,2)-V center ) 2 ) in, H center is the hue range θ H The central value of S center is the saturation range θ S The center value of V center They are the brightness range θ V The central value of I HSV (x,y,0) represents image I HSV The hue component of the pixel (x, y) in the HSV color space; I HSV (x,y,1) represents image I HSV The saturation component of the HSV color space of the pixel (x, y); I HSV (x,y,2) represents image I HSV The brightness component of the HSV color space of the pixel (x, y); σ(·) is the Sigmoid function; w H ,w S ,w V is the learnable weight; 24) For the probability mask P color Normalize and get the color attention weight map A color : Among them, min(P color ) is the probability mask P color The minimum value of max(P color ) is the probability mask P color The maximum value of 25) Color attention weight map A color Perform Gaussian filtering to obtain the color attention weight map A after Gaussian filtering. color ; 3) Obtain the attention-guiding feature map; the specific process is: 31) The RGB image of the rice seedling stage farmland preprocessed in step S2 is passed through a convolution block Conv weed Generate attention weight map A weed ; expressed as: A weed =σ(Conv weed (F shallow )), Among them, σ is the Sigmoid activation function; Conv weed Convolution block, Conv weed It includes convolution layer, maximum pooling layer, and convolution layer in sequence; A weed ∈R H×W×1 ; F shallow The RGB image of the farmland at the seedling stage after preprocessing in step S2; 32)、 The saliency mask attention map M saliency After passing through two convolutional layers in sequence, the feature map F is obtained. saliency ; The color attention weight map A after Gaussian filtering color After passing through two convolutional layers in sequence, the feature map F is obtained. color ; The attention weight map A weed After passing through two convolutional layers in sequence, the feature map F is obtained. weed ; 33) The feature map F saliency , feature map F color , feature map F weed Splicing in the channel dimension, and then passing through the linear layer to obtain the feature map after dimensionality reduction; The feature map after dimensionality reduction and the feature map A output by the first C3k2 module are input into the cross attention module for fusion to obtain the final attention-guided feature map 9. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 8, characterized in that: The specific working process of the first cross-band attention fusion module CBAF is as follows: 1)、 The multispectral images of 450nm and 660nm bands are spliced ​​in the channel dimension, and then input into the seedling state feature extractor 1 for feature extraction to obtain the damaged seedling feature D1; The seedling state feature extractor 1 is SCFE1; SCFE1 includes the first convolution block, the second convolution block, and the third convolution block in sequence; Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a ReLU activation function layer, a convolution layer, a ReLU activation function layer, a convolution layer, and a maximum pooling layer in sequence; 2)、 The multispectral images of 750nm and 840nm bands are spliced ​​in the channel dimension, and then input into the seedling state feature extractor 2 for feature extraction to obtain the floating seedling feature D2; The seedling state feature extractor 2 is SCFE2; SCFE2 includes the first convolution block, the second convolution block, and the third convolution block in sequence; Each convolution block in the first convolution block, the second convolution block, and the third convolution block includes a convolution layer, a ReLU activation function layer, a convolution layer, a ReLU activation function layer, a convolution layer, and a maximum pooling layer in sequence; 3) Cross-attention fusion: The damaged feature D1 is used as the key K, and the feature map B output by the second C3k2 module is used as the query Q and value V respectively to perform cross-attention fusion operation to obtain the feature map The floating feature D2 is used as the key K, and the feature map B output by the second C3k2 module is used as the query Q and value V respectively, and the cross-attention fusion operation is performed to obtain the feature map The calculation formula of the cross-attention fusion operation is as follows: Among them, d k is the dimension of key K; the superscript T indicates transposition; softmax indicates activation function; Attention(Q,K,V) indicates cross-attention fusion operation; 4) The feature map obtained by cross-attention fusion operation and Fusion is performed to obtain the cross-band attention map F CBAF ; The specific process is: Fusion is channel concatenation followed by dimensionality reduction through a convolutional layer: Among them, Conv fuse is the convolutional layer used to process the fused feature map, Concat() is concatenation, and F CBAF is the cross-band attention map; 5) The cross-band attention map F CBAF The Hadamard product is performed with the feature map B output by the second C3k2 module to obtain the features after attention Among them, ⊙ is the Hadamard product.

10. The method for detecting the quality of transplanting posture based on multimodal perception according to claim 9, characterized in that: The deep learning model is trained to obtain a trained deep learning model; the specific process is: (1)、 The RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model, and the first prediction head of the deep learning model outputs the category prediction probability distribution p; Calculate the cross entropy loss for each sample based on the true label y and the predicted probability distribution p output by the first prediction head The calculation formula is as follows: in, C is the number of categories, which is 3; y i is the true label of each sample belonging to the i-th category; p i is the probability p of each sample belonging to the i-th category predicted by the deep learning model i ; i=1, 2 or 3, i represents qualified seedlings, injured seedlings or floating seedlings; (2)、 Step S2: The RGB image and multispectral image of the seedling-stage farmland preprocessed are input into the deep learning model. The first prediction head of the deep learning model outputs the category prediction probability distribution p and calculates the prediction entropy H(p) of each sample. The calculation formula is as follows: Use the Dropout technique to perform T forward propagations and obtain T predicted probability distributions p (1) ,p (2) ,...,p (T) , calculate the prediction variance Var(p) of each sample, the calculation formula is as follows: Among them, Var(·) represents variance calculation; It represents the probability that each sample predicted by the first prediction head of the deep learning model belongs to the i-th category when the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the first time; It indicates that the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the second time, and the probability that each sample belongs to the i-th category predicted by the first prediction head of the deep learning model; represents the probability that each sample predicted by the first prediction head of the deep learning model belongs to the i-th category when the RGB image and multispectral image of the farmland in the seedling stage preprocessed in step S2 are input into the deep learning model for the Tth time; (3) Construct a scoring function for each sample: The scoring function is as follows: Among them, α, β, γ are hyperparameters; (4) Calculate the score S of all samples and select the K% of samples with the highest scores as hard examples based on the scores; (5)、 The weight function w(S) associated with the score of hard examples is defined as follows: w(S)=1+λ×tanh(S) Among them, λ is a hyperparameter that controls the magnitude of weight adjustment; For each hard example, the gradient is calculated using the weighted loss function: in, S(p i ) is the score of the difficult sample calculated using the scoring function; w(S(p i )) is the weight of the hard example; For each non-hard example, the gradient is calculated using a weighted loss function with a weight of 1: Use the back-propagation algorithm to update deep learning model parameters; (6) Train the deep learning model until the weighted loss function Converge and obtain a trained deep learning model.

Citation Information

Cited By

  • Dynamically adaptive vehicle-mounted multispectral farmland weed detection method

    CN121746933A