Melon seedling feature recognition and analysis method based on image multi-dimensional features
By acquiring multimodal information containing two-dimensional images and three-dimensional point cloud data, performing spatial calibration and preprocessing, and combining cross-modal attention mechanism for feature fusion, dynamic fusion feature vectors are generated, and the health status and quantitative growth parameters of seedlings are output. This solves the problem of insufficient robustness of two-dimensional image recognition methods in complex environments, and achieves high accuracy and robustness in seedling recognition.
Patent Information
- Application Number
- CN202511453358.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-10-13
AI Technical Summary
In existing technologies, recognition methods based on two-dimensional images lack robustness in complex seedling cultivation scenarios, making it difficult to capture the three-dimensional spatial morphology of seedlings in real growth environments, resulting in poor recognition performance. Existing fusion strategies are also unable to adapt to the morphological diversity and environmental complexity during the seedling growth process.
By acquiring agricultural robots that autonomously navigate the area, two-dimensional images and three-dimensional point cloud data are simultaneously acquired. Spatial calibration and preprocessing are performed to generate dynamic fusion feature vectors, and the health status and quantitative growth parameters of the target are output.
This invention addresses unresolved technical issues in existing technologies by achieving adaptability to multimodal data. By combining technical means, it provides a method for identifying and analyzing the features of melon seedlings based on multidimensional image features, which can effectively solve the problems mentioned in the background technology.
Smart Images

Figure CN120953749B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent agriculture, and particularly relates to a melon seedling feature recognition and analysis method based on image multi-dimensional features. BACKGROUND
[0002] With the global population growth and the increasing demand for food safety, agriculture is gradually transforming towards intelligence and refinement. The automatic recognition and growth state analysis of melon seedlings have become key technologies for factory seedling and unmanned cultivation systems, and their precision and efficiency directly affect the quality of seedling and the intelligence level of subsequent field management.
[0003] Currently, the related technology mainly relies on a recognition method based on two-dimensional images, which realizes seedling classification and preliminary evaluation by extracting color, texture or morphological features combined with machine learning algorithms. However, two-dimensional images lack depth information and are difficult to capture the three-dimensional spatial form of seedlings in the real growth environment, and are easily disturbed by occlusion, light changes and plant overlap in complex seedling scenes, with insufficient recognition robustness.
[0004] Although some research has introduced multi-modal data fusion methods to try to combine two-dimensional images and three-dimensional point clouds to improve recognition results, existing fusion strategies mostly use feature splicing or fixed weight summation methods, which fail to fully consider the complementary structure and dynamic association between different modal features, resulting in discriminative features being submerged and making it difficult to adapt to the morphological diversity and environmental complexity of seedling growth. SUMMARY
[0005] To overcome the shortcomings in the background art, the embodiments of the present application provide a melon seedling feature recognition and analysis method based on image multi-dimensional features, which can effectively solve the problems involved in the above background art.
[0006] The purpose of the present application can be achieved by the following technical solutions: a melon seedling feature recognition and analysis method based on image multi-dimensional features, comprising: autonomously cruising the melon crop seedling area by an agricultural robot, and synchronously acquiring multi-modal data containing two-dimensional image information and three-dimensional point cloud information of the area.
[0007] The multi-modal data is spatially calibrated and preprocessed to extract two-dimensional semantic features and three-dimensional geometric features of the target seedling, and the current growth stage of the target seedling is identified based on the extracted features.
[0008] According to the current growth stage of the target seedling, the fusion weights of the two-dimensional semantic features and the three-dimensional geometric features are dynamically adjusted, and a cross-modal attention mechanism is used for feature fusion to generate a dynamic fusion feature vector.
[0009] Input the dynamic fusion feature vector into a task recognition model to output the health state and quantitative growth parameters of the target seedling.
[0010] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects: (1) The present application breaks through the limitations of traditional calibration methods in unstructured agricultural scenes by performing adaptive spatial calibration and data deep preprocessing on multi-modal data, effectively avoiding information pollution caused by information misplacement, and ensuring the quality of original data.
[0011] (2) The present application creatively proposes a dynamic multi-modal feature fusion method based on growth stage perception, which can adaptively adjust the fusion weights of two-dimensional visual features and three-dimensional geometric features according to the real-time growth state of seedlings, and establish deep correlation between semantic and geometric information through cross-modal interactive attention mechanism, effectively avoiding feature redundancy and key information dilution caused by traditional static fusion strategies, and significantly improving the robustness and measurement accuracy of seedling recognition in complex seedling raising scenes.
[0012] (3) The present application outputs the health state and quantitative growth parameters of the target seedling in parallel through the task recognition model, not only realizes comprehensive and multi-dimensional evaluation, but also further improves the accuracy and robustness of overall recognition analysis by utilizing the mutual promotion effect between tasks. BRIEF DESCRIPTION OF DRAWINGS
[0013] The present application will be further described with the aid of the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present application, and other drawings can be obtained by those of ordinary skill in the art without creative labor on the basis of the following drawings.
[0014] Figure 1 The present application is a method step flowchart.
[0015] Figure 2 The present application is a bidirectional cross-modal attention interaction logic diagram.
[0016] Figure 3 The present application is a logic diagram for obtaining and outputting the health state and quantitative growth parameters of the target seedling. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0018] REFERENCE Figure 1As shown, the application provides a melon seedling feature recognition analysis method based on image multi-dimensional features, comprising: S11. An agricultural robot autonomously patrols a melon crop seedling area, and synchronously acquires multi-modal data containing two-dimensional image information and three-dimensional point cloud information of the area.
[0019] It should be noted that the synchronous acquisition of two-dimensional image information and three-dimensional point cloud information is achieved through a hardware trigger synchronization mechanism: the main controller of the agricultural robot generates a synchronization trigger signal and simultaneously sends it to the image acquisition component and the depth perception component.
[0020] The image acquisition component and the depth perception component carried by the agricultural robot start exposure and data collection at the same time when receiving the trigger signal, ensuring that the two-dimensional image and the three-dimensional point cloud data have strict time synchronization.
[0021] S12. Space calibration and preprocessing are performed on the multi-modal data to extract two-dimensional semantic features and three-dimensional geometric features of the target seedling, and the current growth stage of the target seedling is identified based on the extracted features.
[0022] In a preferred embodiment of the application, the space calibration includes: identifying a pre-set artificial marker in the seedling area from the synchronously acquired two-dimensional image information and three-dimensional point cloud information.
[0023] It should be noted that the artificial marker has texture features and geometric features that are significantly different from the seedling area, for identification in two-dimensional images and three-dimensional point clouds.
[0024] The key points and their descriptors of the artificial marker in the two-dimensional image and the corresponding key points and their three-dimensional geometric descriptors in the three-dimensional point cloud are extracted respectively.
[0025] It should be noted that the extraction of the key points and their descriptors of the artificial marker in the two-dimensional image is obtained by performing a scale-invariant feature transform algorithm on the two-dimensional image.
[0026] The extraction of the corresponding key points and their three-dimensional geometric descriptors in the three-dimensional point cloud is achieved through the following steps: first, for each sampling point in the three-dimensional point cloud, perform principal component analysis in its pre-set K neighborhood, K being the number of neighborhood points, which needs to be adjusted adaptively according to the point cloud density, solve the principal component direction by decomposing the spatial coordinate matrix of the neighborhood points, and then calculate the normal vector of the point.
[0027] At the same time, a local covariance matrix is constructed based on the spatial distribution of the neighborhood points, and the positions where the geometric morphology in the local area of the point cloud changes significantly are identified by analyzing the relative numerical differences of the eigenvalues of the covariance matrix, such as the eigenvalue variance ratio or the maximum to minimum eigenvalue ratio, to serve as three-dimensional key points, and generate three-dimensional geometric descriptors in combination with the normal vector direction.
[0028] Perform keypoint matching based on the two-dimensional and three-dimensional descriptors to solve the coordinate transformation parameters between the image acquisition component and the depth perception component carried by the agricultural robot.
[0029] Iteratively optimize the coordinate transformation parameters so that the re-projection error of the three-dimensional point cloud projected into the two-dimensional image coordinate system is less than a set threshold.
[0030] It should be noted that the above-mentioned coordinate transformation parameters include a rotation matrix and a translation vector, and the solving process is specifically to find potential matching pairs by calculating the Euclidean distance between two-dimensional descriptors and three-dimensional geometric descriptors, to identify and eliminate false matches by randomly selecting a minimum data set, to retrieve the best matching pairs between the two descriptors, and to solve the initial coordinate transformation parameters including the rotation matrix and the translation vector by least squares fitting according to each best matching pair.
[0031] Project the three-dimensional point cloud into the image coordinate system according to the current transformation parameters to establish a correspondence between each three-dimensional point and its nearest key point on the two-dimensional image.
[0032] Based on the established correspondence, a least squares objective function of the distance between the point pairs is constructed.
[0033] The optimal rotation matrix and translation vector are solved by singular value decomposition to update the transformation parameters.
[0034] Repeat the iteration until the re-projection error converges below the set threshold, and finally obtain the accurate coordinate transformation parameters.
[0035] In a preferred embodiment of the present application, the preprocessing includes two-dimensional image preprocessing: local contrast enhancement and noise filtering processing on the two-dimensional image.
[0036] It should be noted that the local contrast enhancement is specifically implemented by dividing the image into multiple small regions and performing histogram equalization on each region.
[0037] The noise filtering processing is specifically implemented by calculating the current pixel value by weighted average of all pixel blocks in the image, wherein the weight depends on the similarity between the pixel blocks.
[0038] Based on the visual feature difference between the seedlings and the background area in the image, the processed image is classified at the pixel level, the pixels are divided into background class and seedling class, and all seedling class pixels are aggregated to generate a binary mask representing the melon seedling area.
[0039] Based on the binary mask, color space statistical features, texture distribution features and two-dimensional morphological geometric features are extracted from the defined seedling area to obtain the two-dimensional semantic features of the target seedling.
[0040] It should be noted that the color space statistical features include but are not limited to the RGB, HSV and Lab color space mean, standard deviation, skewness and kurtosis of the pixels in the melon seedling region.
[0041] The texture distribution features specifically refer to the contrast, correlation index, energy and homogeneity index extracted by using the gray level co-occurrence matrix.
[0042] The two-dimensional morphological geometric features specifically refer to the area and perimeter obtained by calculating the number of pixels and the number of boundary pixels of the segmentation mask, the aspect ratio calculated by fitting the minimum circumscribed rectangle or ellipse, the compactness calculated by the ratio of the area to the perimeter, the compactness is used to reflect the stretching or gathering degree of the shape of the seedling, and the Fourier descriptor is used to describe the contour complexity of the leaf of the melon seedling.
[0043] In a preferred embodiment of the present application, the preprocessing further includes three-dimensional point cloud preprocessing: outlier filtering and data simplification processing are performed on the three-dimensional point cloud to obtain purified point cloud data.
[0044] It should be noted that the outlier filtering is specifically performed by calculating the average distance in the K-neighborhood of each point cloud, and if the average distance of a certain point exceeds the preset standard deviation multiple of the global average distance, the point is marked as an outlier and filtered out.
[0045] The data simplification processing specifically adopts a voxel downsampling method: the point cloud space is divided into three-dimensional cubic voxels, and the centroid of all points in each voxel is used to replace all points in the voxel, thereby greatly reducing the number of points while maintaining the overall geometric structure of the point cloud, and the voxel size can be exemplarily set as a cubic grid shape of 3mm x 3mm x 3mm.
[0046] The connected seedling regions in the purified point cloud data are segmented into independent single plant point clouds.
[0047] For each independent seedling point cloud, its normal vector and curvature are calculated as local geometric attributes, and based thereon, the plant height, volume, stem thickness, leaf angle and spread of the single plant are extracted to jointly constitute the three-dimensional geometric features of the target seedling.
[0048] It should be noted that the plant height of the single plant is obtained by calculating the difference between the maximum coordinate value in the Z-axis direction of the point cloud and the Z-axis coordinate value of the ground reference plane.
[0049] After voxelizing the point cloud, the local density correction is performed on the structures such as leaf wrinkles and stem bending in combination with the curvature information, the point cloud voxel space contained or involved by the single plant is counted, and the planning volume of each voxel unit is accumulated, so as to estimate the plant volume.
[0050] The local point cloud slice is intercepted in the stem region, a circular cross section is fitted, and the average or median of the cross section radius is taken as the estimated value of the stem thickness, in which the normal vector is used to determine the axial direction of the stem, and the curvature is used to identify the boundary between the stem and the leaves, branches and other parts, so as to avoid non-stem points from participating in fitting.
[0051] The angle between the normal vector direction of the leaf point cloud region and the horizontal plane is taken as the leaf inclination.
[0052] The leaf point cloud is projected onto the horizontal plane, and the minimum circumscribed rectangle or convex hull of the projected point set is extracted, and the maximum length and width thereof are taken as the leaf spread indicators.
[0053] In a preferred embodiment of the present application, the connected seedling regions in the purified point cloud data are divided into independent single plant point clouds, including: performing Euclidean cluster analysis on the purified point cloud data, and dividing the seedling point cloud into multiple independent connected regions based on a point cloud spatial distance threshold.
[0054] The orientation correction based on principal component analysis is performed on each independent connected region, so that the main axis is aligned with the vertical direction.
[0055] Based on the corrected point cloud region, the minimum bounding box size is calculated, and the abnormal regions with sizes not meeting the expected range of single seedling are filtered out, and the effective single plant point cloud data is retained.
[0056] The embodiments of the present application break through the limitations of traditional calibration methods in unstructured agricultural scenes by performing adaptive spatial calibration and data depth preprocessing on multi-modal data, not only effectively avoiding information pollution caused by information misplacement, but also ensuring the quality of original data.
[0057] In a preferred embodiment of the present application, the current growth stage of the target seedling includes: determining the standard growth cycle of the target seedling based on the type to which the target seedling belongs, and dividing the standard growth cycle into multiple continuous growth stages.
[0058] Based on the historical planting data, the typical patterns of each growth stage in three-dimensional geometry and two-dimensional semantic features are established.
[0059] It should be noted that the establishment of the typical patterns of each growth stage in three-dimensional geometry and two-dimensional semantic features is implemented as follows: according to the planting records of the target seedling of the same category of crops in the seedling area, multi-modal data samples containing the melon seedling in different growth stages are collected, and each sample has an accurate growth stage label.
[0060] The two-dimensional semantic features and three-dimensional geometric features are extracted from each sample data.
[0061] For each growth stage, two-dimensional features and three-dimensional features of all samples in the stage are respectively subjected to cluster analysis to obtain cluster centers, wherein the cluster analysis specifically adopts existing cluster algorithms such as K-Means or K-Medoids algorithm.
[0062] The cluster centers are defined as typical modes of two-dimensional and three-dimensional features of the growth stage.
[0063] According to the currently extracted three-dimensional geometric features, a dominant interval of distribution of the single-plant population in the sequence growth stage in the seedling area is determined.
[0064] It should be noted that the dominant interval determination process is: according to the three-dimensional geometric features of the target seedling currently extracted, the matching degrees of the typical modes of the three-dimensional geometric features of each single plant in the seedling area with respect to each growth stage are analyzed, the analysis method can exemplarily adopt a cosine similarity calculation formula, the growth stage with the largest matching degree is determined as the tendency stage of the single plant, the tendency stages of all single plants are counted, a number of growth periods with a high frequency of occurrence and continuity are integrated to preliminarily determine the dominant interval, the high requirement specifically refers to that the proportion of the number of single plants in the same tendency stage to all plants in the seedling area reaches a preset percentage, the preset percentage is flexibly set in combination with the sample size of the seedling area and the fineness of the growth stage division, specifically follows the principle of positive correlation with the sample size and negative correlation with the fineness of the growth stage division, and can be experimentally calibrated by a method developer.
[0065] The growth stages in the dominant interval are marked as candidate stages.
[0066] According to the matching degrees of the currently extracted two-dimensional semantic features and the typical modes of each candidate stage, the candidate stage with the highest matching degree is determined as the current growth stage.
[0067] S13. According to the current growth stage of the target seedling, the fusion weights of the two-dimensional semantic features and the three-dimensional geometric features are dynamically adjusted, and a cross-modal attention mechanism is used for feature fusion to generate a dynamic fusion feature vector.
[0068] In a preferred embodiment of the present application, the dynamic adjustment of the fusion weight is realized by the following way: the corresponding relationship rules of different growth stages and fusion weights are pre-stored, and the rules are established based on different dependence degrees of the two-dimensional semantic features and the three-dimensional geometric features for the health state recognition task in different growth stages.
[0069] It should be noted that the establishment of the corresponding relationship rules is realized by the following steps: (1) collecting a multi-modal data set of target seedlings containing multiple health states in different growth stages, each sample being labeled with an accurate growth stage label and a health state label.
[0070] (2) Using only two-dimensional features, only three-dimensional features, and different weight combinations of fusion features, respectively, to verify the health state recognition accuracy.
[0071] (3) For each growth stage, select the weight combination that optimizes the recognition performance.
[0072] (4) The optimal weight combination corresponding to each growth stage is summarized as a clear mapping relationship to form a final weight distribution rule library.
[0073] The correspondence rule is established on the basis that the morphological structure and physiological characteristics of seedlings are different at different growth stages, resulting in changes in the importance of different modalities of features in health state recognition.
[0074] The following is an exemplary weight distribution table, only for reference: Table 1 is an example of weight distribution of two-dimensional semantic features and three-dimensional geometric features for some growth stages
[0075] Table 1
[0076]
[0077] According to the identified current growth stage, the corresponding fusion weight is assigned by calling the correspondence rule.
[0078] If necessary, the dynamic weight adjustment optimizes the quality of the final fusion feature vector, which is used for both health state classification and growth parameter regression. The dynamic adjustment of the fusion weight based on the health state rather than the quantitative growth parameter identification task is because the health state recognition task has the strongest selectivity and the most obvious dynamic change in feature modalities, and best reflects the advantages of the adaptive fusion mechanism of the present application. In addition, a fusion feature that more accurately reflects the true health status of seedlings must also contain more abundant and accurate structural information. Therefore, the fusion strategy oriented to health state optimization not only improves the accuracy of health state recognition, but also objectively provides higher quality feature input for growth parameter regression, thereby indirectly ensuring the accuracy of growth parameter measurement
[0079] Referring to Figure 2 In a preferred embodiment of the present application, the cross-modal attention mechanism is used for feature fusion, which includes: performing self-attention calculation on the two-dimensional semantic features and three-dimensional geometric features to obtain enhanced two-dimensional semantic features and three-dimensional geometric features, and performing bidirectional cross-modal attention interaction, including: a. The enhanced features of the first modality are used as query vectors, and the enhanced features of the second modality are used as key vectors and value vectors.
[0080] b. Calculate the dot product of the query vector and the key vector to obtain the initial similarity between each feature element of the first modality and each feature element of the second modality, and build a cross-modal similarity matrix based on this.
[0081] c. Normalize the similarity matrix to generate an attention weight matrix representing the attention degree of the feature elements of the first modality to the feature elements of the second modality.
[0082] d. Weighted sum the attention weight matrix and the value vector to generate the second modality attention feature integrated with the context of the first modality.
[0083] e. Exchange the roles of the first modality and the second modality, repeat steps a-d to generate the first modality attention feature integrated with the context of the second modality.
[0084] Based on the results of the bidirectional cross-modal attention interaction, combine the dynamically adjusted fusion weight to generate the dynamic fusion feature vector.
[0085] It should be noted that the above self-attention calculation implementation process and the bidirectional cross-modal attention interaction implementation process have logical similarities, which are implemented as follows: map the feature vectors of each modality to query, key and value vectors respectively.
[0086] Calculate the similarity of the query and key vectors and normalize to obtain the attention weight matrix.
[0087] Weighted fusion of the attention weight matrix and the value vector to generate an enhanced feature vector.
[0088] In the self-attention mechanism: the attention weight matrix represents the importance between the feature elements within the modality. For example, in two-dimensional semantic features, the color features of a certain leaf area may be highly related to the texture features of another area, and the weight matrix will give higher importance to this internal association.
[0089] In the cross-modal attention mechanism: the attention weight matrix represents the importance of the feature elements between the modalities. For example, a geometric key point in a three-dimensional point cloud will calculate its association strength with all two-dimensional image feature points, so as to determine which two-dimensional texture information is most important for explaining the semantics of the three-dimensional point.
[0090] Therefore, the importance is not pre-set, but the association strength learned adaptively through query-key similarity calculation. The weight matrix is a direct and unique numerical representation of this importance measure, which guides the focus on the most relevant information in the feature fusion process.
[0091] The embodiment of the application creatively proposes a dynamic multi-modal feature fusion mode based on growth stage perception, which can adaptively adjust the fusion weight of two-dimensional visual features and three-dimensional geometric features according to the real-time growth state of seedlings, and establish a deep correlation between semantic and geometric information through a cross-modal interactive attention mechanism, effectively avoiding feature redundancy and key information dilution caused by traditional static fusion strategies, and significantly improving the robustness and measurement accuracy of seedling recognition in complex seedling raising scenes.
[0092] Referring to Figure 3 As shown in S14, the dynamic fusion feature vector is input into a task recognition model, and the health state and quantitative growth parameters of the target seedling are output.
[0093] In a preferred embodiment of the application, the task recognition model is constructed by: obtaining a training sample set, each sample including the dynamic fusion feature vector, and corresponding health state label and at least one quantitative growth parameter true value.
[0094] It should be noted that the health state label is in the form of multi-label encoding, indicating health, water deficiency, nutrient deficiency, disease, etc.
[0095] The quantitative growth parameter true value is the real measurement value of parameters such as plant height, leaf area, stem diameter, etc. obtained by manual measurement or high-precision instruments.
[0096] The sample set is divided into a training set, a validation set and a test set according to a predetermined ratio, which can be exemplarily 7:2:1, to ensure the representativeness of data distribution.
[0097] A deep neural network is constructed, which includes a shared bottom network and at least two output branches for health state classification and quantitative growth parameter regression, respectively.
[0098] The deep neural network is trained by jointly optimizing the classification loss and the regression loss to obtain the task recognition model.
[0099] It should be noted that the quantitative growth parameter is different from the quantified growth parameter of the two-dimensional semantic feature and the three-dimensional geometric feature, including but not limited to the following categories: crown width, projected leaf area, stem diameter, leaf curl index, plant compactness, biomass estimate, the quantitative growth parameter is simultaneously considered by the task recognition model to make more accurate and robust predictions than a single modality by considering the complex nonlinear relationship between its three-dimensional morphology and two-dimensional growth.
[0100] In a preferred embodiment of the application, it further includes: during the autonomous cruise of the agricultural robot, the spatial position coordinates corresponding to each set of synchronously collected multi-modal data are recorded in real time by the positioning system thereof.
[0101] The health state and the quantitative growth parameter output by the task recognition model for each group of data are associated and bound with the spatial position coordinates recorded when the corresponding data are collected.
[0102] Based on the recognition results of all the bound position information, a spatial interpolation algorithm is used to generate a health state distribution map and a quantitative growth parameter distribution map covering the whole seedling raising area, and the maps are visually output.
[0103] The embodiments of the present application parallelly output the health state and the quantitative growth parameter of the target seedling through the task recognition model, not only realize comprehensive and multi-dimensional evaluation, but also further improve the accuracy and robustness of the overall recognition analysis by utilizing the mutual promotion effect between tasks.
[0104] It needs to be particularly pointed out that, in order to accurately understand the technical solutions of the present application, the following terms are clarified: In the present application, the melon crop seedling raising area is a macro geographical range concept, which refers to the whole operation area including seedlings, soil, seedbed and the like. The seedling area is a target object area accurately extracted from the above macro area through image segmentation or point cloud segmentation technology, and specifically refers to the pixel or point cloud set occupied by the melon seedling itself.
[0105] In other words, the seedling area is a foreground target separated from the seedling raising area. The core of the present application is to process the whole data of the seedling raising area, and finally accurately locate, analyze and recognize the characteristics and state of the seedling area.
[0106] The above content is only an example and description of the structure of the present application. Those skilled in the art can make various modifications or supplements or use similar ways to replace the described specific embodiments, as long as they do not deviate from the structure of the present application or exceed the scope defined by the present application, which shall belong to the protection scope of the present application.
Claims
1. A method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features, characterized in that, include: The agricultural robot autonomously navigates the seedling area of cucurbit crops, simultaneously acquiring multimodal data containing two-dimensional image information and three-dimensional point cloud information of the area; Spatial calibration and preprocessing are performed on the multimodal data to extract the two-dimensional semantic features and three-dimensional geometric features of the target seedling, and the current growth stage of the target seedling is identified based on the extracted features. Based on the current growth stage of the target seedling, the fusion weights of the two-dimensional semantic features and the three-dimensional geometric features are dynamically adjusted, and a cross-modal attention mechanism is used for feature fusion to generate a dynamic fused feature vector. The dynamically fused feature vector is input into the task recognition model, which outputs the health status and quantitative growth parameters of the target seedling. The identification of the current growth stage of the target seedling includes: determining its standard growth cycle based on the type of the target seedling, and dividing the standard growth cycle into multiple consecutive growth stages; establishing typical patterns of each growth stage in terms of three-dimensional geometry and two-dimensional semantic features based on historical planting data; determining the dominant interval of the distribution of individual plant populations in the sequenced growth stages within the seedling area based on the currently extracted three-dimensional geometric features; marking the growth stages within the dominant intervals as candidate stages; and determining the candidate stage with the highest matching degree as the current growth stage based on the matching degree between the currently extracted two-dimensional semantic features and the typical patterns of each candidate stage.
2. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The space calibration includes: Artificial markers pre-set in the seedling area are identified from the synchronously acquired two-dimensional image information and three-dimensional point cloud information. The key points and their descriptors of the artificial markers in the two-dimensional image, as well as the corresponding key points and their three-dimensional geometric descriptors in the three-dimensional point cloud, are extracted respectively. Based on the two-dimensional and three-dimensional descriptors, key point matching is performed to solve the coordinate transformation parameters between the image acquisition component and the depth perception component carried by the agricultural robot. The coordinate transformation parameters are iteratively optimized to ensure that the reprojection error of the 3D point cloud onto the 2D image coordinate system is less than a set threshold.
3. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The preprocessing includes two-dimensional image preprocessing: The two-dimensional image is subjected to local contrast enhancement and noise filtering. Based on the visual feature differences between seedling and background regions in the image, the processed image is classified at the pixel level, and the pixels are divided into background and seedling categories. All seedling category pixels are aggregated to generate a binary mask representing the melon seedling region. Based on the binary mask, color space statistical features, texture distribution features, and two-dimensional morphological geometric features are extracted from the defined seedling area to obtain the two-dimensional semantic features of the target seedling.
4. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The preprocessing also includes 3D point cloud preprocessing: The three-dimensional point cloud is subjected to outlier filtering and data simplification processing to obtain purified point cloud data; The connected seedling regions in the purified point cloud data are divided into independent single-plant point clouds. For each individual seedling point cloud, its normal vector and curvature are calculated as local geometric attributes. Based on these, the plant height, volume, stem thickness, leaf inclination angle, and spread of a single plant are extracted, which together constitute the three-dimensional geometric features of the target seedling.
5. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 4, characterized in that, The purified point cloud data is divided into connected seedling regions into independent single-plant point clouds, including: Euclidean clustering analysis was performed on the purified point cloud data to segment the seedling point cloud into multiple independent connected regions based on the point cloud spatial distance threshold. For each independent connected region, perform orientation correction based on principal component analysis to align its principal axis with the vertical direction; Based on the corrected point cloud region, its minimum bounding box size is calculated, and abnormal regions whose sizes do not conform to the expected range of a single seedling are filtered out, while retaining the valid single-seedling point cloud data.
6. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The dynamic adjustment of the fusion weights is achieved through the following methods: There are pre-stored rules for the correspondence between different growth stages and fusion weights. These rules are established based on the different dependencies of the health status recognition task on two-dimensional semantic features and three-dimensional geometric features at different growth stages. Based on the identified current growth stage, the corresponding relationship rule is invoked to assign the corresponding fusion weight.
7. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The feature fusion using a cross-modal attention mechanism includes: Self-attention calculations are performed on the two-dimensional semantic features and the three-dimensional geometric features respectively to obtain enhanced two-dimensional semantic features and three-dimensional geometric features, and bidirectional cross-modal attention interaction is performed, including: a. Use the enhanced features of the first modality as the query vector, and the enhanced features of the second modality as the key vector and value vector; b. By calculating the dot product of the query vector and the key vector, the initial similarity between each feature element of the first modality and each feature element of the second modality is obtained, and a cross-modal similarity matrix is constructed accordingly; c. Normalize the similarity matrix to generate an attention weight matrix that represents the degree of attention paid by the first modality feature elements to the second modality feature elements; d. The attention weight matrix and the value vector are weighted and summed to generate a second modality attention feature integrated into the first modality context; e. Swap the roles of the first modality and the second modality, repeat step ad, and generate the first modality attention features integrated into the context of the second modality; Based on the results of the bidirectional cross-modal attention interaction, and combined with dynamically adjusted fusion weights, the dynamic fusion feature vector is generated.
8. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, The task recognition model construction includes: Obtain a training sample set, each sample including the dynamic fusion feature vector, the corresponding health status label, and at least one quantitative growth parameter ground value; Construct a deep neural network, which includes a shared underlying network and at least two output branches, used for health status classification and quantitative growth parameter regression, respectively; The deep neural network is trained by jointly optimizing the classification loss and regression loss to obtain the task recognition model.
9. The method for feature recognition and analysis of cucurbit seedlings based on multidimensional image features according to claim 1, characterized in that, Also includes: During the autonomous navigation process of the agricultural robot, its positioning system records the spatial coordinates corresponding to each set of synchronously collected multimodal data in real time. The health status and quantitative growth parameters output by the task recognition model for each set of data are associated and bound with the spatial location coordinates recorded during the corresponding data collection. Based on the identification results of all bound location information, a spatial interpolation algorithm is used to generate a health status distribution map and a quantitative growth parameter distribution map covering the entire seedling area, and then visualizes the output.
Citation Information
Patent Citations
Noisy plant point cloud semantic segmentation method and system based on self-attention feature fusion
CN116311218A
Melon seedling phenotypic characteristic measurement and digital grading system based on image processing
CN118823448A