Tea bud screening method based on multi-modal information fusion
By combining multimodal information fusion and environmental perception adaptive mechanisms with image, point cloud and spectral data, the problems of low efficiency and low accuracy in traditional tea bud screening are solved, and high-precision intelligent tea bud picking is achieved.
Patent Information
- Application Number
- CN202510993367.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional tea bud screening relies on manual labor, which is inefficient and costly. Single machine vision tea bud recognition has low accuracy and is difficult to achieve accurate screening in complex environments.
A multimodal information fusion method is adopted, which combines industrial cameras, lidar and near-infrared sensors to collect images, point clouds and spectral data of tea tree areas. Data preprocessing is performed through illumination compensation network, structure-texture separation network, point cloud filtering network and graph convolutional network to extract multiple features and perform decision-level fusion to drive the robotic arm to perform tea bud picking.
It improves the accuracy of tea bud identification and screening efficiency in complex environments, and realizes high-precision intelligent harvesting.
Smart Images

Figure CN120877046A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer and image recognition, and specifically to a method for screening tea buds based on multimodal information fusion. Background Technology
[0002] Traditional tea bud screening relies heavily on manual picking and sorting, which is inefficient, labor-intensive, and costly. With the development of machine vision technology, machine vision-based tea bud screening methods have gradually been applied. However, single machine vision image information is greatly affected by factors such as lighting and the tea growing environment, resulting in low accuracy in tea bud identification and susceptibility to misjudgment. When encountering interference from similar leaves or bud occlusion, relying solely on image features such as color and shape is insufficient to accurately distinguish tea buds from non-bud parts, leading to insufficient screening precision and failing to meet the demands of high-quality tea processing. Existing technologies have attempted to improve the problem by optimizing image processing algorithms, but due to the limited information source, accurate screening remains difficult to achieve in complex environments. Summary of the Invention
[0003] To address the aforementioned shortcomings in existing technologies, this invention provides a tea bud screening method based on multimodal information fusion. It further introduces an environmental awareness adaptive mechanism, dynamically estimating the confidence level of each modality based on the image signal-to-noise ratio (SNR), point cloud occlusion degree (Occ), and spectral signal intensity, and automatically adjusting the fusion weight of each modality, thereby improving the robustness of recognition and the success rate of picking in complex environments.
[0004] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A tea bud screening method based on multimodal information fusion includes the following steps: S1. Simultaneously acquire images, point clouds, and spectral heterogeneous data of the tea tree area using industrial cameras, lidar, and near-infrared sensors; S2. Based on the heterogeneous data, an illumination compensation network is used to optimize image brightness and a structure-texture separation network is used to perform image denoising. At the same time, a point cloud filtering network is applied to perform density adaptive denoising and non-rigid registration on the point cloud data, and a graph convolutional network is constructed to model key bands of the spectral data. S3. Based on the preprocessed data, extract the color, texture and shape features of the image, the spatial structure features of the point cloud and the physiological features of the spectrum, and generate a joint feature vector through a feature-level fusion network. S4. Based on the joint feature vector, a comprehensive decision score is generated by dynamically weighting the confidence scores of each modality using a decision-level fusion network. S5. When the comprehensive decision score exceeds the adaptive threshold, the robotic arm is driven to perform tea bud picking action based on the spatial coordinates reconstructed from the point cloud.
[0005] Furthermore, step S2 specifically includes the following steps: S21. Apply an illumination compensation network to the image data to perform exposure correction. The output function is: in, To adjust the image at the pixel level Color channels The brightness value on; This represents the brightness value of the corresponding pixel and channel in the original input image; The overall brightness offset factor is the output of the prediction subnetwork; These are the color channel weight parameters learned during the training process; S22. Apply a structure-texture separation network to the image data for denoising. The output function is: in, For the optimized image after denoising, Output for the main structure, This is the response adjustment factor for the texture residual branch at each location. Output for texture branches; S23. Apply a point cloud filtering network to the point cloud data to perform density adaptive noise filtering, calculate the density variation coefficient, and enable weighted median filtering when the density variation function is greater than the set threshold. S24. Apply a non-rigid registration algorithm to the point cloud data, and optimize the objective function as follows: in, Let be a rotation matrix. It is a translation vector. It is a non-rigid micro-transformation phase. This represents the non-rigid micro-deformation term at each point. The regularization coefficient is... It is a point The deformation gradient measures the change of deformation in space; S25. Apply a graph convolutional network to the spectral data to construct a sparse spectral graph while preserving the boundaries.
[0006] Furthermore, the training loss function of the structure-texture separation network in S22 is: In the formula, To improve the pixel absolute error between the image and the high-quality reference image, For structural similarity constraints, For image gradient preservation, These are the corresponding hyperparameters.
[0007] Furthermore, the total loss function of the graph convolutional network in S25 is: In the formula, For basic classification cross-entropy loss, For spectral distribution preservation term, For spectral structure smoothing term, is the regularization coefficient.
[0008] Furthermore, step S3 specifically includes the following steps: S31. Apply the Mobile-Former network to the image data output by S2 to extract color features, texture features, and shape features; S32. Extract the geometric center coordinates, principal axis direction, and surface curvature features from the point cloud data output by S2. S33. Apply a graph convolutional network to the spectral data output from S2 to extract the absorption peak characteristics of water, chlorophyll, and amino acids. S34. The features of each modality are concatenated into a joint feature vector through a feature-level fusion network.
[0009] Furthermore, step S4 specifically includes the following steps: S41. Based on the joint feature vector of S3, a decision-level fusion network is applied to calculate the image confidence S. img Spectral matching degree S spec and point cloud pickability score S pc ; S42. Based on the results calculated in S41, a comprehensive decision score is generated through dynamic weighting.
[0010] Furthermore, S5 specifically includes the following steps: S51. The target's 3D pose reconstructed based on the S2 point cloud network is mapped to the robotic arm's coordinate system; S52. Use the inverse kinematics solver to plan the robot arm path; S53. Compensate for trajectory errors at a frequency of 1kHz via servo control; S54 triggers the pneumatic gripper to perform pressure feedback gripping.
[0011] The present invention has the following beneficial effects: This invention aims to provide a tea bud screening method based on multimodal information fusion and environmental perception adaptive mechanism, which improves the identification accuracy and screening efficiency of tea buds in complex environments, and realizes a high-precision, scalable intelligent harvesting application. Attached Figure Description
[0012] Figure 1This is a schematic diagram of the tea bud screening method based on multimodal information fusion according to the present invention. Detailed Implementation
[0013] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0014] A tea bud screening method based on multimodal information fusion, such as Figure 1 As shown, it includes the following steps: S1. Simultaneously acquire images, point clouds, and spectral heterogeneous data of the tea tree area using industrial cameras, lidar, and near-infrared sensors; In this embodiment, the machine vision camera deployment involves using high-resolution industrial cameras for image acquisition to ensure clear capture of the dynamic details of tea trees and buds. For camera placement, a matrix deployment scheme is adopted, taking into account the actual terrain features and planting density of the tea garden. Camera arrays are installed above the tea trees, ensuring at least 20% overlap in the field of view between adjacent cameras to avoid blind spots. In key recognition areas, a dual / multi-macro camera joint registration mechanism is introduced. Two or more macro cameras are configured in key areas of the harvesting operation to simultaneously capture the same tea bud target from different angles. By performing stereo matching and registration on multi-view images, a disparity map is constructed and depth information is restored, accurately reconstructing the three-dimensional spatial position of the tea bud, significantly improving recognition and positioning accuracy. This stereo vision solution is particularly suitable for real-world scenarios where tea buds have complex structures, small sizes, and are often obstructed, providing high-quality visual guidance for the precise harvesting of the subsequent robotic arm end effector. In addition, all cameras are equipped with an automatically adjustable optical filter module that can dynamically switch filter modes according to the ambient light intensity, suppressing overexposure in strong light conditions and enhancing image brightness in low light environments, thereby further improving image clarity and recognition robustness.
[0015] LiDAR equipment installation: A mechanical rotating LiDAR is used, with a scanning angle range of 360° and a vertical field of view of no less than ±15°, ensuring complete coverage of the tea tree's growing space. The LiDAR's ranging accuracy reaches the millimeter level, and the point cloud density is set to no less than 500 points / square meter. During installation, the LiDAR is fixed on a height-adjustable bracket. The bracket height can be electrically adjusted according to the tea tree's growth height to ensure accurate 3D point cloud data is acquired at different growth stages.
[0016] Near-infrared sensor configuration: A near-infrared sensor with a wavelength range of 760-1100nm was selected. This wavelength band is sensitive to the water content of plant tissues and can effectively reflect the physiological differences between tea buds and mature leaves. The sensor's sampling frequency was set to 10Hz to synchronize with the data acquisition frequency of the machine vision camera and LiDAR. Near-infrared sensors were evenly distributed between the rows of tea bushes, with each sensor having a detection angle of 60° to ensure that all parts of the tea bush could be effectively detected.
[0017] Time synchronization mechanism established: A high-precision GPS clock synchronization module is used to provide a unified time reference for machine vision cameras, LiDAR, and near-infrared sensors. At the data acquisition end, a precise timestamp is added to each frame of data, with a timestamp resolution down to the microsecond level. Simultaneously, a synchronization calibration procedure based on the Network Time Protocol (NTP) is designed to automatically calibrate the time of each device every 10 minutes, ensuring that the data time alignment error does not exceed 1 millisecond.
[0018] S2. Based on the heterogeneous data, an illumination compensation network is used to optimize image brightness and a structure-texture separation network is used to perform image denoising. At the same time, a point cloud filtering network is applied to perform density adaptive denoising and non-rigid registration on the point cloud data, and a graph convolutional network is constructed to model key bands of the spectral data.
[0019] The illumination compensation network comprises two sub-modules: an illumination deviation prediction sub-network and a conditional exposure control sub-network. The prediction sub-network constructs an illumination shift model based on regional brightness statistics; the control sub-network dynamically enhances the RGB three-channel brightness through a channel attention mechanism and incorporates seasonal and time-of-day factors for illumination inference, improving the system's image adaptability under complex lighting conditions and optimizing the input image. The image is segmented into multiple non-overlapping local regions, which are then input into a self-built lightweight convolutional network to predict the overall brightness shift of the image. This is used to measure the degree of brightness deviation between the current image and the standard tea garden exposure template. The network structure includes 3 convolutional layers (3×3 kernel size), ReLU activation and global average pooling units, and finally outputs a single scalar value as a conditional factor for exposure control.
[0020] In addition to the image's grayscale mean and brightness histogram distribution, the input features also incorporate regional contrast grading statistics and color channel shift features, forming a multi-dimensional perceptual vector. The network training objective is to minimize the following brightness deviation loss function: in The standard exposure is collected from an agricultural sample library (which can be set to the average brightness of the reference image set). To improve robustness, an exposure smoothing regularization term is introduced during network training to constrain the continuity of predicted values, ensuring consistency of predictions across time and region dimensions.
[0021] In addition, this module references prior knowledge of agricultural light (such as light variation time curves and statistical distribution of tea garden samples) and introduces seasonal and time-period calibration coefficients during the inference stage. The final brightness shift prediction is: To achieve an exposure perception output that better matches the actual lighting conditions in tea gardens.
[0022] This embodiment specifically includes the following steps: S21. Apply an illumination compensation network to the image data to perform exposure correction. This module adjusts the brightness enhancement level of each color channel in the image based on the predicted brightness deviation. Its network structure consists of three convolutional layers, a channel attention module (SE-Block), and a channel transformation module. Its function is to perform non-equilibrium weighted adjustment of the brightness of the RGB three channels according to the current exposure state to enhance the overall visibility and color balance of the image. Its output function is: in, To adjust the image at the pixel level Color channels The brightness value on; This represents the brightness value of the corresponding pixel and channel in the original input image; The overall brightness offset factor is the output of the prediction subnetwork; These are the color channel weight parameters learned during training. They measure the sensitivity of different channels to exposure adjustments, ensuring that the image colors remain balanced during the adjustment process. S22. Apply a structure-texture separation network to the image data for denoising. The output function is: in, For the optimized image after denoising, Output for the main structure, This is the response adjustment factor for the texture residual branch at each location. Output for texture branches; The adjusted image, as a preprocessed result, will be directly used by subsequent feature extraction networks (such as Mobile-Former) for the identification and classification of tea bud regions. This adjustment process, while maintaining structural clarity, effectively enhances the contrast in low-light areas and improves the color separation of leaf regions, which is of great significance for improving the accuracy of target edge extraction and discrimination.
[0023] The network training employs an unsupervised contrastive learning approach, and the following joint loss function is designed. : Brightness expectation loss makes the output image brightness approach the mean exposure value of the agricultural sample set; Edge preservation loss constrains gradient map changes to prevent enhancement from causing structural ambiguity; Color channel balance loss prevents color cast caused by exposure adjustments; , , These are hyperparameter weights, with typical values of 10, 5, and 8.
[0024] In its implementation, the AGENet backbone network employs a four-layer residual convolutional structure, with each layer containing a 3×3 convolutional kernel, ReLU activation, and batch normalization. The conditional modulation module constructs a weight tensor based on a channel attention mechanism to perform biased enhancement on the brightness channel of each pixel.
[0025] To address the issues of drastic changes in signal-to-noise ratio (SNR) and easy loss of detailed structure in natural tea garden scenes, a structure-texture self-separating image denoising network (ST-DeNet) is proposed. This network suppresses background noise in the original image while accurately preserving the edge contours and detailed texture information of tea buds. Based on a dual-branch structure, the network extracts the structural backbone and local texture residuals of the image separately, and introduces a dynamic residual gating mechanism to achieve region-aware noise control.
[0026] A multi-scale residual coding module is employed to extract the main contour and large structural block information of the image, focusing on modeling the shape continuity between the background region and the tea bud boundary. This branch uses a combination of layer-by-layer pyramidal convolution and dilated convolution with increasing dilation rate to enhance cross-scale structure perception. The structural backbone branch is responsible for extracting large structural information from the input image, focusing on the shape continuity of the tea bud and its surrounding background. The network structure of this branch includes multi-scale convolutional layers, which can effectively model different spatial scales of the image. The detailed network architecture is as follows: Input layer: Input the original image (size, for example, 224×224×3).
[0027] Convolutional layers: First convolutional layer: Uses a 3x3 convolution kernel, stride 1, padding "same", output feature map size is 224×224×64.
[0028] The second convolutional layer uses a 3x3 kernel with a stride of 2 and padding with "same". The output feature map size is 112×112×128.
[0029] The third convolutional layer uses a 3x3 kernel with a stride of 2 and padding with "same". The output feature map size is 56×56×256.
[0030] Pyramid convolution module: Uses pyramid convolution structures in intermediate layers (e.g., progressively increasing the receptive field) to enhance cross-scale information fusion capabilities.
[0031] Dilated Convolution Module: This module uses dilated convolutions to expand the receptive field without increasing computational cost. It is particularly important for modeling long-range dependencies.
[0032] Fully connected layer: A fully connected layer is used to flatten the feature map output by the convolutional layer, and the output is a vector of fixed dimensions (e.g., 1024 dimensions).
[0033] Output layer: Outputs feature vectors containing large structural information, which are then used for subsequent fusion with other branches.
[0034] 2) Texture Residual Branch Focusing on modeling high-frequency texture information (such as leaf veins, fine thorns, jagged edges, etc.), this branch employs shallow convolutional stacking combined with an edge-aware enhancement module to capture subtle detail changes in the image. This branch preserves detail frequency features through independent paths, preventing over-smoothing in the main branch. The texture residual branch is responsible for extracting detail information from the image, such as leaf veins, edges, and fine textures. This branch models these high-frequency detail information through independent paths, avoiding over-smoothing. The detailed network architecture is as follows: Input layer: Input the original image (size 224x224x3).
[0035] Convolutional layers: First convolutional layer: Uses a 3x3 convolution kernel, stride 1, padding "same", output feature map size is 224×224×64.
[0036] The second convolutional layer uses a 3x3 kernel with a stride of 1 and padding with "same". The output feature map size is 224×224×128.
[0037] The third convolutional layer uses a 3x3 kernel with a stride of 1 and padding with "same". The output feature map size is 224×224×256.
[0038] Texture enhancement module: A detail enhancement module is introduced to capture detailed features and generate texture feature maps by using specific convolution kernels.
[0039] Local features are extracted using deep convolutional networks, and the perception of detailed textures is enhanced by increasing the number of layers.
[0040] Residual learning module: Residual connections are used to avoid information loss during training. Residual learning allows the network to focus on learning subtle differences.
[0041] By combining shallow convolutions with detail enhancement modules, subtle changes in image detail can be captured.
[0042] Fully connected layer: Flatten the features output by the convolutional layer and input them into the fully connected layer, outputting a vector of size 1024.
[0043] Output layer: Outputs texture feature vectors with detailed information, which are then fused with other branches.
[0044] 3) Adaptive Residual Modulation Unit (ARMU) Introduce a residual gating mechanism based on SNR awareness: in This is the response adjustment factor for the texture residual branch at each location. and For empirical parameters (such as) = 1.0, = 0.2), ε is the signal-to-noise ratio at the current point, and ε is a small constant to prevent the denominator from being zero. Through this mechanism, weighted enhancement can be achieved for low signal-to-noise ratio regions, and weighted suppression can be achieved for high signal-to-noise ratio regions, thereby improving the network's denoising selectivity in complex regions.
[0045] 4) Fusion and Reconstruction Module Finally, the backbone structure output is weighted and fused with the adjusted texture residual information to output an optimized image with clear structure, rich details, and suppressed noise. : in Output for the main structure, Output for texture branches.
[0046] The network training employs a multi-loss joint optimization strategy: : Enhance the absolute pixel error between the image and the high-quality reference image; Structural similarity constraints prevent the noise reduction process from disrupting the image's spatial layout; Image gradient preservation term, which enhances edge perception in textured regions; Typical hyperparameter settings are α=1, β=1, δ=0.5.
[0047] This method demonstrates excellent image denoising robustness and detail preservation in tea garden images with strong natural light variations and significant local occlusion.
[0048] S23. Apply a point cloud filtering network to the point cloud data to perform density adaptive noise filtering, calculate the density variation coefficient, and enable weighted median filtering when the density variation function is greater than the set threshold. To improve the robustness of lidar in processing point cloud data in natural tea garden environments, especially in noise reduction and multi-view stitching effects in scenes with dense foliage and severe spatial occlusion, a density adaptive residual fusion point cloud fine filtering algorithm (DARF-PC) is proposed. Through density change detection, geometric consistency constraints, and residual-guided registration methods, efficient noise reduction and non-rigid registration of point cloud data are achieved.
[0049] 1) Point cloud anomaly detection and density adaptive filtering First, the original point cloud is divided into local voxel mesh regions (typically 5cm × 5cm × 5cm), and the point cloud density variation coefficient is calculated within each local region. : in, The standard deviation of the number of points within that unit voxel. This represents the mean number of points within a unit voxel. If the density variation in a certain region is significantly greater than that in the surrounding area (i.e., ...), then... > ,like If the value is less than 1.5, it is considered an abnormal region and enters the adaptive filtering stage.
[0050] For regions with normal density, mean filtering with a medium window is applied; for regions with abrupt density changes, a weighted median filter is introduced, with the kernel size equal to... It is proportional to the spatial density, thus dynamically adjusting the filtering intensity.
[0051] 2) Local geometric consistency constraints eliminate jump points Based on the filtering, the system further determines whether point p is a structural transition point. It calculates the average normal angle between p and its k nearest neighbors. : in, It is the normal vector of point p, representing the orientation of the surface at that point. If is the normal vector of the i-th nearest neighbor, then > If the angle is 25°, it is considered a geometric change region and can be eliminated or retained depending on the task objective (for example, the standard can be relaxed to retain the harvesting area).
[0052] 3) Non-rigid registration: Residual-driven improved ICP (NR-ICP) Traditional ICP algorithms are not very accurate in non-rigid point cloud deformation caused by branch swaying or occlusion. Therefore, this system proposes a residual-driven non-rigid iterative registration method (NR-ICP) to optimize the multi-view registration effect of point clouds.
[0053] The specific steps are as follows: In the initial stage, a rigid initial fit is performed based on spindle alignment; Constructing the local residual field: For a point p in the source point cloud, the error between it and the nearest point q(p) in the target point cloud is defined as: in, : Represents the residual of point p between the source point cloud and the target point cloud, that is, the difference between point p in the source point cloud and the nearest point in the target point cloud; The point corresponding to point p in the source point cloud comes from the target point cloud; : Rotation matrix, representing the rotation transformation of the point cloud; Translation vector, representing the translation transformation of the point cloud; Handling of residual regularization terms and registration errors.
[0054] S24. Apply a non-rigid registration algorithm to the point cloud data. During registration, the goal of the regularization term is to guide the optimization process, ensuring that it not only minimizes the direct residual between the source and target point clouds but also smooths and optimizes non-rigid deformation regions within the point clouds. Specifically, the regularization term is adjusted and controlled in the following ways: Construction of regular terms In the objective function, the residual regularization term is constructed considering the geometric deformation of the point cloud, and a local deformation field is introduced to balance the registration between deformed and non-deformed regions. The final objective function is: in, Let be a rotation matrix. It is a translation vector. It is a non-rigid micro-transformation phase. This represents the non-rigid micro-deformation term at each point. The regularization coefficient is... It is a point The deformation gradient measures the change of deformation in space; How to balance deformation and residual By minimizing This optimization objective aims to reduce the error between the source and target point clouds. However, point clouds may exhibit local deformation, especially in complex scenes (such as when branches or tea buds obscure the view). Therefore, a regularization term is introduced. The objective function aims not only to reduce the residuals but also to ensure the deformation term. Maintain smoothness and avoid excessive local deformation to prevent the accumulation of errors during registration.
[0055] iii. How to control the smoothness of deformation Regularization term The role of regularization in the objective function is to avoid excessive deformation by controlling the magnitude of the deformation gradient. This means that if the deformation gradient at a point is large, it indicates that the point cloud in that region has undergone drastic changes, which may lead to inaccurate registration. By introducing the regularization term, the optimization process tends to reduce such drastic deformation, thereby smoothing the transition and making the point cloud registration result more accurate.
[0056] Combination with the final registration target The ultimate goal is to achieve precise alignment between the source and target point clouds. Regularization ensures that when calculating the point cloud residuals, optimization is not only applied to positional errors but also considers the smoothness of the overall deformation, thus avoiding unnatural abrupt changes—a crucial aspect of non-rigid point cloud registration. Specifically, regularization ensures the stability of the point cloud registration process, resulting in a smooth transition over a large area, rather than localized distortion.
[0057] Dynamic adjustments during the optimization process The smoothness λ of the regularization term is a key parameter that controls the tolerance for deformation during optimization. In practical applications, λ can be adjusted through an iterative process to adapt to the point cloud characteristics in different scenarios. For example, in scenarios with significant occlusion or uneven object surfaces, a smaller λ value allows for greater local deformation; while in scenarios with more regular structures, a larger λ value helps maintain higher smoothness and avoid error accumulation. S25. Apply a graph convolutional network to the spectral data to construct a sparse spectral graph while preserving the boundaries.
[0058] In near-infrared spectral data processing, to improve identification specificity and the efficiency of spectral information utilization, key bands were systematically selected in the 760–1100 nm range based on expert experience, sample grouping statistical analysis, and significance testing (p < 0.05). Representative bands that most significantly reflected the responses of tea buds to moisture, chlorophyll, and amino acids were identified, such as 780 nm (main absorption peak for moisture), 860 nm (carotenoid activity region), 940 nm (chlorophyll content variation window), and 1050 nm (amino acid band). This constructed a set of physiologically relevant node features. Each selected band was used as a node input to the GCN model to capture the spectral structural differences of vegetation targets.
[0059] In terms of spectral structure construction, a collaborative strategy of similarity filtering and redundancy suppression based on spectral weight construction is proposed. This is achieved by calculating the cosine similarity matrix between normalized spectral response vectors. And introduce spectral weighting coefficients (Determined based on the contribution rate of the principal components of the band), retain Edges exceeding a dynamic threshold δ (e.g., δ = μ + 1.5σ, where μ is the mean and σ is the standard deviation) are removed to eliminate redundant weakly correlated connections, compress the graph structure, and improve embedding focus. Regarding the loss function design, in addition to the basic classification cross-entropy loss... In addition, the following two regularization designs are further introduced: 1) Spectral distribution preservation term This prevents high-dimensional feature collapse and cluster drift by minimizing the squared distance between the node vector within a category and its category center. The formula is as follows: in Embed for the i-th node, It is the center vector of its category; 2) Spectral structure smoothing term Encouraging coherent changes in features between adjacent nodes in the spectral graph enhances the interpretability of the embedding vector within the spectral domain, defined as: Where E represents the set of edges in the final sparse graph.
[0060] The final total loss function is expressed as: Here, α and β are regularization coefficients (e.g., α = 0.3, β = 0.5), and the optimal configuration can be obtained through a validation set grid search. Through the multi-dimensional refinement at the levels of node selection, spectral structure modeling, and regularization term design, this spectral module not only enhances the model's ability to perceive subtle spectral differences in complex farmland backgrounds but also significantly improves the robustness and interpretability of the tea bud screening system in actual deployment.
[0061] S3. Based on the preprocessed data, extract the color, texture and shape features of the image, the spatial structure features of the point cloud and the physiological features of the spectrum, and generate a joint feature vector through a feature-level fusion network. This embodiment specifically includes the following steps: S31. Apply the Mobile-Former network to the image data output by S2 to extract color features, texture features, and shape features; The core of the feature-level fusion module lies in deep encoding and unified representation of key features from different sensory sources to construct a multimodal high-dimensional fusion vector with discriminative power and robustness.
[0062] In the image modality, the preprocessed image is first enhanced in brightness and edge details using the AGENet network, and then denoised with high quality using the ST-DeNet structure—texture self-separation mechanism. For feature extraction, the Mobile-Former network is used to obtain shallow and global semantic features of the image, modeling three types of visual descriptors: color, texture, and shape. Specifically: Color features are modeled using the RGB channel mean and standard deviation to represent the differences between illumination and color. Texture features are calculated using the gray-level co-occurrence matrix to determine energy, contrast, and entropy, reflecting leaf surface details. Shape features are obtained by using edge detection and contour analysis to acquire boundary attributes such as area, perimeter, and eccentricity; In addition, a dual / multi-macro camera deployment mechanism is introduced to achieve stereo parallax estimation of the target area, obtain local depth maps, and calculate three-dimensional geometric features such as concavity and convexity, edge gradient and elevation changes, thereby improving the target separability in complex occlusion scenes.
[0063] S32. Extract the geometric center coordinates, principal axis direction, and surface curvature features from the point cloud data output by S2. After point cloud modalities are processed by DARF-PC filtering and NR-ICP non-rigid registration, spatial structural features such as geometric center coordinates (x, y, z), principal axis direction, point density, and surface curvature of each candidate target region are extracted. This information is used to express the three-dimensional configuration, posture, and distribution characteristics of tea buds.
[0064] S33. Apply a graph convolutional network to the spectral data output from S2 to extract the absorption peak characteristics of water, chlorophyll, and amino acids. In terms of spectral modes, a sparse spectral structure is constructed based on key bands, and the response modes between bands are extracted using a graph convolutional network (GCN). The resulting dense spectral vectors (16-32 dimensions) reflecting the absorption peak characteristics of water, chlorophyll and amino acids are embedded to form a spectral description of the physiological state of tea buds.
[0065] S34. The features of each modality are concatenated into a joint feature vector through a feature-level fusion network.
[0066] After all modal features are standardized by z-score (mean 0, standard deviation 1), they are sequentially concatenated in vector space to form a fused vector. This serves as the input for the subsequent fusion and discrimination module. This fusion representation fully integrates multi-dimensional features of structure, spectrum, and texture, exhibiting higher recognition robustness and tea bud differentiation ability in complex natural environments.
[0067] S4. Based on the joint feature vector, a comprehensive decision score is generated by dynamically weighting the confidence scores of each modality using a decision-level fusion network. The decision-level fusion module is responsible for integrating the recognition and scoring results of three modalities: image, point cloud, and spectrum, to generate a final judgment score to determine whether the target area is a "pickable tea bud".
[0068] In the image modality, the system uses the image preprocessed with AGENet + ST-DeNet as input to the Mobile-Former recognition network. It extracts regional structural features through shallow convolutions and models global image relationships using a lightweight Transformer module, outputting a recognition confidence score. This reflects the degree of matching between visual features and the target tea bud.
[0069] In the spectral modality, the GCN network performs graph convolutional embedding learning on the spectrum constructed from key bands to generate high-dimensional spectral response vectors, and the MLP classifier outputs the corresponding probability scores. This score measures the similarity between the spectral characteristics of the current region and the physiological characteristics of the tea bud sample.
[0070] In the point cloud modality, based on the 3D data preprocessed by DARF-PC + NR-ICP, the system comprehensively evaluates parameters such as the geometric stability, occlusion degree, principal axis orientation, and gripper path reachability of the target region. A lightweight pluckability estimation network is used to generate an executable score. This score measures whether the current target is ready for harvesting.
[0071] Final judgment score The scores of the three modalities are determined according to their respective weighting coefficients. , Weighted fusion: The weights of each modality are not fixed but dynamically estimated by the real-time perception module. The system introduces an adaptive adjustment mechanism based on signal-to-noise ratio (SNR) and occlusion degree (Occ), estimating the confidence factors of the three modalities according to the current image sharpness, point cloud density integrity, and spectral signal stability. Then normalize to the final fusion weight: This mechanism ensures that modalities with higher information quality have higher decision weights in complex lighting or strong occlusion environments, thereby improving the robustness and fault tolerance of the overall judgment.
[0072] When the fusion score If the area exceeds the system's set threshold θ (which is automatically learned by maximizing the F1-score on the validation set), the system determines that the area is a target for picking tea buds and calls the execution module in real time to issue a grasping command, driving the robotic arm to perform precise picking operations.
[0073] S5. When the comprehensive decision score exceeds the adaptive threshold, the robotic arm is driven to perform tea bud picking action based on the spatial coordinates reconstructed from the point cloud.
[0074] When the decision score output by the multimodal fusion model When the adaptive threshold θ is exceeded (which is automatically learned by maximizing the F1-score in the validation set), the system will enter the tea bud picking execution stage. First, based on the target position coordinates given by the fusion judgment module and combined with the point cloud depth information provided by the LiDAR, the system reconstructs the precise position and orientation of the tea bud in three-dimensional space. This spatial pose is then mapped to the robot world coordinate system as a reference target for the robotic arm's movements.
[0075] After receiving the target pose, the robotic arm control system initiates the motion planning module and uses an inverse kinematics (IK) solver to calculate the optimal arrival path for the six-DOF robotic arm. During this process, the system comprehensively considers the joint limits of the robotic arm, gripper orientation constraints, and environmental obstacles within the workspace, generating a smooth and collision-free path through collision detection and trajectory optimization algorithms (such as RRT* or CHOMP).
[0076] During execution, the servo control system monitors the joint position and end-effector deviation at a frequency of 1 kHz to achieve closed-loop control compensation. Encoder feedback corrects end-effector trajectory errors in real time, ensuring the gripper is ultimately aligned with the target position to meet harvesting accuracy requirements. Once the gripper approaches the tea bud target, the pneumatic drive device begins the gripping action, and pressure sensor feedback ensures appropriate clamping force to avoid damaging the tea bud.
[0077] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0078] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0079] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0080] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0081] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A method for screening tea buds based on multimodal information fusion, characterized in that, Includes the following steps: S1. Simultaneously acquire images, point clouds, and spectral heterogeneous data of the tea tree area using industrial cameras, lidar, and near-infrared sensors; S2. Based on the heterogeneous data, an illumination compensation network is used to optimize image brightness and a structure-texture separation network is used to perform image denoising. At the same time, a point cloud filtering network is applied to perform density adaptive denoising and non-rigid registration on the point cloud data, and a graph convolutional network is constructed to model key bands of the spectral data. S3. Based on the preprocessed data, extract the color, texture and shape features of the image, the spatial structure features of the point cloud and the physiological features of the spectrum, and generate a joint feature vector through a feature-level fusion network. S4. Based on the joint feature vector, a comprehensive decision score is generated by dynamically weighting the confidence scores of each modality using a decision-level fusion network. S5. When the comprehensive decision score exceeds the adaptive threshold, the robotic arm is driven to perform tea bud picking action based on the spatial coordinates reconstructed from the point cloud.
2. The tea bud screening method based on multimodal information fusion according to claim 1, characterized in that, S2 specifically includes the following steps: S21. Apply an illumination compensation network to the image data to perform exposure correction. The output function is: in, To adjust the image at the pixel level Color channels The brightness value on; This represents the brightness value of the corresponding pixel and channel in the original input image; The overall brightness offset factor is the output of the prediction subnetwork; These are the color channel weight parameters learned during the training process; S22. Apply a structure-texture separation network to the image data for denoising. The output function is: in, For the optimized image after denoising, Output for the main structure, This is the response adjustment factor for the texture residual branch at each location. Output for texture branches; S23. Apply a point cloud filtering network to the point cloud data to perform density adaptive noise filtering, calculate the density variation coefficient, and enable weighted median filtering when the density variation function is greater than the set threshold. S24. Apply a non-rigid registration algorithm to the point cloud data, and optimize the objective function as follows: in, Let be a rotation matrix. It is a translation vector. It is a non-rigid micro-transformation phase. This represents the non-rigid micro-deformation term at each point. The regularization coefficient is... It is a point The deformation gradient measures the change of deformation in space; S25. Apply a graph convolutional network to the spectral data to construct a sparse spectral graph while preserving the boundaries.
3. The tea bud screening method based on multimodal information fusion according to claim 2, characterized in that, The training loss function of the structure-texture separation network in S22 is: In the formula, To improve the pixel absolute error between the image and the high-quality reference image, For structural similarity constraints, For image gradient preservation, These are the corresponding hyperparameters.
4. The tea bud screening method based on multimodal information fusion according to claim 2, characterized in that, The total loss function of the convolutional network in S25 is: In the formula, For basic classification cross-entropy loss, For spectral distribution preservation term, For spectral structure smoothing term, is the regularization coefficient.
5. The tea bud screening method based on multimodal information fusion according to claim 1, characterized in that, S3 specifically includes the following steps: S31. Apply the Mobile-Former network to the image data output by S2 to extract color features, texture features, and shape features; S32. Extract the geometric center coordinates, principal axis direction, and surface curvature features from the point cloud data output by S2. S33. Apply a graph convolutional network to the spectral data output from S2 to extract the absorption peak characteristics of water, chlorophyll, and amino acids. S34. The features of each modality are concatenated into a joint feature vector through a feature-level fusion network.
6. The tea bud screening method based on multimodal information fusion according to claim 1, characterized in that, S4 specifically includes the following steps: S41. Based on the joint feature vector of S3, a decision-level fusion network is applied to calculate the image confidence S. img Spectral matching degree S spec and point cloud pickability score S pc ; S42. Based on the results calculated in S41, a comprehensive decision score is generated through dynamic weighting.
7. The tea bud screening method based on multimodal information fusion according to claim 1, characterized in that, S5 specifically includes the following steps: S51. The target's 3D pose reconstructed based on the S2 point cloud network is mapped to the robotic arm's coordinate system; S52. Use the inverse kinematics solver to plan the robot arm path; S53. Compensate for trajectory errors at a frequency of 1kHz via servo control; S54 triggers the pneumatic gripper to perform pressure feedback gripping.
Citation Information
Cited By
Greenhouse plant product harvesting monitoring method and system based on knowledge graph
CN121121502A
Intelligent tea leaf sorting method based on multi-modal perception and adaptive learning
CN122200640A
A tea intelligent sorting method based on multi-modal perception and adaptive learning
CN122200640B