Spatial heterogeneous visual positioning system and positioning method of pineapple picking robot
The spatial heterogeneous visual positioning system of the pineapple picking robot based on multi-source information fusion solves the problems of positioning accuracy and fruit defect detection in complex terrain, realizes efficient and stable pineapple picking, and improves picking quality and efficiency.
Patent Information
- Application Number
- CN202510860889.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing pineapple picking robots lack positioning accuracy and robustness in complex terrain, making it difficult to achieve efficient and non-destructive detection of internal defects in the fruit, and the picking efficiency and quality are low.
A spatial heterogeneous visual positioning system for pineapple picking robots adopts multi-source information fusion, combining vision, inertial measurement, multispectral and structured light data, and performs posture prediction through extended Kalman filtering and nonlinear optimization methods. It also uses a multimodal deep learning model for defect detection and combines a heuristic search algorithm to plan the motion path.
It achieves high-precision spatial positioning and posture estimation in complex terrain, improves the stability and real-time performance of picking, improves the accuracy of internal defect detection of fruits and picking quality, optimizes the picking path, and reduces the risk of mechanical damage.
Smart Images

Figure CN120593732A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of picking robot positioning, and in particular to a spatially heterogeneous visual positioning system and a positioning method for a pineapple picking robot. Background Art
[0002] With the acceleration of agricultural modernization, the application of intelligent harvesting robots in the field of fruit and vegetable picking is becoming increasingly widespread. Especially in the automatic harvesting tasks under complex terrain conditions, the accuracy and robustness of the robot positioning and control system have become a key technical bottleneck. Traditional harvesting robots mostly rely on single visual or inertial sensors for environmental perception and posture estimation. They are unable to meet the requirements of high-precision spatial positioning in uneven terrain such as terraces and slopes. This leads to increased positioning errors of the robotic arm and unstable picking movements, which affects picking efficiency and fruit quality. In addition, the ability to detect internal defects of the fruit is insufficient, making it impossible to achieve intelligent graded picking. This is obviously insufficient to improve picking quality and reduce subsequent processing costs.
[0003] Prior art publication CN119772891A discloses a spatially heterogeneous visual positioning system for a tomato-picking robot, involving the field of deep learning. The system includes the following steps: Step S1: calibrating the positioning of the tomato-picking robot; Step S2: acquiring tomato images in a non-enclosed environment using a fixed camera; Step S3: acquiring spatial coordinate information of the fruit within the fixed camera's field of view; Step S4: determining the reachable space of the picking robot and moving the picking robot; Step S5: performing secondary spatial perception of the target using a palm-mounted camera to further acquire spatial coordinate information in a relatively enclosed environment; Step S6: performing approximate sphere fitting to obtain the optimal picking target; and Step S7: determining the number of tomatoes and estimating the picking pose. While this system can improve positioning accuracy, it relies on multiple precise calibrations and multi-stage visual perception, resulting in low real-time and robustness. Furthermore, it lacks multi-sensor fusion, insufficient intelligent fusion, and insufficient dynamic adaptability, making it difficult to ensure high-precision positioning and stable picking in complex environments, limiting operational efficiency and environmental adaptability.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a spatially heterogeneous visual positioning system and positioning method for a pineapple picking robot to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A spatially heterogeneous visual positioning system for a pineapple picking robot, specifically comprising:
[0008] a data acquisition module, electrically connected to the posture prediction module and the path planning module, for collecting visual data, inertial measurement data, position data, and multispectral-structured light data of the robot in complex terrain, and sending the data to the posture prediction module;
[0009] A posture prediction module, electrically connected to the manipulator control module, configured to predict the robot's posture based on the visual data and the inertial measurement data, and to send the posture prediction result to the manipulator control module;
[0010] a defect detection module, the defect detection module being electrically connected to the robotic arm control module and configured to detect and perform preliminary classification of internal defects of the pineapple fruit based on the multispectral-structured light data, and to send the classification results to the robotic arm control module;
[0011] A robotic arm control module, which is used to dynamically adjust the picking strategy and the robot's robotic arm movements based on the posture prediction results and the classification results;
[0012] A path planning module, electrically connected to the drive control module, for generating a motion path for the robot in complex terrain based on the robot's positioning data and preset three-dimensional environmental data, and sending the generated path to the drive control module;
[0013] A drive control module is used to control the robot to move according to the received motion path.
[0014] Preferably, the visual data is an RGB-D image of the robot's position; the inertial measurement data includes the speed, acceleration, and angular velocity of the robot during movement; and the multispectral-structured light data includes a multispectral image and a structured light depth map, and covers near-infrared and short-wave infrared bands.
[0015] Preferably, the posture prediction module estimates the robot posture by combining extended Kalman filtering and nonlinear optimization method when fusing visual data and inertial measurement data. The specific logic is as follows:
[0016] Temporal synchronization and spatial calibration of the visual data and inertial measurement data collected by the robot;
[0017] The robot's posture is predicted based on the extended Kalman filter and updated in combination with visual data;
[0018] Based on nonlinear optimization method, the robot's historical posture is globally optimized to reduce the cumulative error;
[0019] The final output is the real-time position and posture prediction of the robot in complex terrain.
[0020] Preferably, the posture prediction of the robot is defined in the following vector form:
[0021]
[0022] In the formula represents the robot's posture vector at time t, p(t), v(t), q(t) represent the robot's position, velocity, and posture quaternion at time t, respectively. a (t), b g (t) represents the bias of acceleration and angular velocity respectively;
[0023] When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing the cost function, which is expressed as follows:
[0024]
[0025] Where r(IMU,t) represents the IMU pre-integration residual at time t, r(vision,t) represents the visual measurement residual at time t, ∑IMU,t and ∑vision,t represent the covariance matrices of the IMU pre-integration residual and the visual measurement residual at time t, respectively. represents the weighted Euclidean norm.
[0026] Preferably, the defect detection module has a built-in multimodal deep learning model for detecting internal defects of pineapple fruits. The training logic of the multimodal deep learning model is:
[0027] Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the categories of the internal defects were annotated.
[0028] Construct a fusion feature vector of multispectral-structured light, perform weighted processing on it, and then input it into a multimodal deep learning model for training;
[0029] Using cross entropy as the loss function, with the goal of minimizing the loss function, the model parameters of the multimodal deep learning model are repeatedly iterated until the optimization is completed;
[0030] in:
[0031] The expression of the fused feature vector is:
[0032]
[0033] In the formula Represents the fusion feature vector of the pixel at the coordinate (x, y), I(λ i ,x,y) indicates the wavelength is λ i The multispectral reflectance of the pixel at coordinate (x, y) is represented by D(x, y), the structured light depth value of the pixel at coordinate (x, y), k(x, y) represents the curvature feature of the pixel at coordinate (x, y), x and y represent the horizontal and vertical coordinates of the pixel in the image, the subscript i represents the index of the selected wavelength, and N is the total number of selected wavelengths;
[0034] When weighting the fused feature vector, the self-attention mechanism is introduced to extract the contextual latent state features of each pixel point, and the dynamic weight of each vector element is calculated based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as:
[0035]
[0036] Where w ∈ (x,y) represents the dynamic weight of the ∈th vector element in the fusion feature vector, h(x,y) represents the contextual hidden state feature of the pixel at coordinate (x,y), ∈ represents the index of the vector element, and f qz (·) represents the weight generating function;
[0037] Then use the dynamic weight to perform weighted processing on the fusion feature vector, and the calculation method is:
[0038]
[0039] In the formula represents the fused feature vector after weighted processing, ⊙ represents the element-by-element product;
[0040] The expression of the cross entropy loss function is:
[0041]
[0042] Where θ represents the model parameters to be trained, y m 、 They represent the true label and model prediction label of the internal defect category of the sample, the subscript m represents the index of the sample, and λ represents the regularization hyperparameter. represents the spatial gradient of the weights.
[0043] Preferably, the logic for primary classification of pineapple fruits using the trained multimodal deep learning model is:
[0044] The multispectral-structured light data of pineapple fruits collected in real time are fused to obtain the fusion feature vector of multispectral-structured light.
[0045] The fused feature vector is weighted and input into the multimodal deep learning model to output the corresponding internal defect category and defect probability;
[0046] The product of the severity of the internal defect category and the defect probability is used as the defect score, and the quality of pineapple fruit is graded according to the defect score.
[0047] Preferably, the logic of the robotic arm control module dynamically adjusting the picking strategy and the robotic arm motion according to the posture prediction result and the classification result is:
[0048] Determine the position and posture of the robotic arm base based on the real-time position and posture prediction of the robot;
[0049] The robot arm's picking path is planned based on the quality grade of the pineapple fruit, with priority given to picking pineapples with higher quality grades.
[0050] After each picking is completed, the picking order is dynamically adjusted and the picking path of the robotic arm is updated until all pineapple fruits in the area are picked.
[0051] Preferably, the planning logic of the robot motion path is:
[0052] Identify the robot's traversable area during movement based on the robot's position data and a preset terrain model;
[0053] Plan the shortest path from the current area to the next area based on a heuristic search algorithm, avoiding obstacles and steep slopes;
[0054] Generate continuous motion trajectories until all pineapple fruits in the area are picked.
[0055] A spatially heterogeneous visual positioning method for a pineapple picking robot is provided. The positioning method is applicable to the above-mentioned positioning system and specifically comprises the following steps:
[0056] S1: Collects multimodal visual data and inertial measurement data of the complex terrain environment in which the robot is located, and simultaneously collects multispectral and structured light data of the target fruit;
[0057] S2: The robot achieves spatial positioning and posture prediction through a visual-inertial fusion algorithm. It also uses multispectral and structured light data to detect and classify internal fruit defects in real time based on a multimodal deep learning model.
[0058] S3: Control the picking path of the robotic arm based on the robot's position, posture, and pineapple fruit quality level to complete the picking of pineapples in the area where the robot is located;
[0059] S4: The robot's motion path is planned according to the robot's position data and a preset terrain model, and steps S1 to S3 are repeated in each area until the pineapple fruits in all areas are picked.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] The present invention realizes high-precision spatial positioning and posture estimation of robots in complex terrain environments by comprehensively utilizing multi-source information such as vision, inertial measurement, multispectral and structured light, significantly improving positioning stability and real-time performance; through the embedded multimodal deep learning model, efficient non-destructive detection and intelligent classification of internal defects of pineapple fruits are realized, supporting the dynamic adjustment of picking strategies of the robotic arm, giving priority to picking fruits with better quality, and improving picking quality and efficiency; at the same time, combined with the preset three-dimensional terrain model and heuristic search algorithm, the robot's motion path is effectively planned, ensuring the safe and stable movement of the robot in complex environments and reducing the risk of mechanical damage. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 This is a schematic diagram of the overall module structure of the present invention;
[0063] Figure 2 It is a schematic diagram of the overall process of the present invention;
[0064] Figure 3 Parameter comparison diagram of different solutions of the present invention. DETAILED DESCRIPTION
[0065] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.
[0066] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0067] Example:
[0068] See also Figures 1 to 3 , the present invention provides a technical solution:
[0069] A spatial heterogeneous visual positioning system for a pineapple picking robot specifically comprises: a data acquisition module, a posture prediction module, a defect detection module, a robotic arm control module, a path planning module, and a drive control module.
[0070] The data acquisition module is electrically connected to the posture prediction module and the path planning module. It collects visual data, inertial measurement data, position data, and multispectral-structured light data of the pineapple fruit while the robot is in complex terrain, and sends it to the posture prediction module. The visual data consists of an RGB-D image of the robot's position; the inertial measurement data includes the robot's velocity, acceleration, and angular velocity during movement; and the multispectral-structured light data includes a multispectral image and structured light depth map, covering the near-infrared and short-wave infrared bands.
[0071] The data acquisition module can be composed of a variety of hardware structures. For example, the RealSense D455 series RGB-D camera can be used to detect visual data, the BMI 160 series inertial measurement unit can be used to detect inertial measurement data, the NEO-M8N series GPS positioning module can be used to obtain the robot's position data, and the RedEdge-MX series multispectral camera and the RealSense L515 series structured light sensor can be used to jointly detect the multispectral-structured light data of pineapple fruit.
[0072] The data acquisition module is responsible for acquiring multimodal data of the robot's operating environment, including color and depth images, robot motion inertial information, real-time geographic location, and multispectral and structured light information of pineapple fruits. By combining multi-source data, it improves the accuracy and robustness of environmental perception and provides accurate raw data for subsequent data processing and decision-making.
[0073] The posture prediction module is electrically connected to the robotic arm control module, and is used to predict the robot's posture based on the visual data and inertial measurement data, and send the posture prediction result to the robotic arm control module.
[0074] The posture prediction module can use the AGX Xavier series of industrial-grade embedded computing platforms and can also be configured with Xilinx Zynq Ultrascale series FPGA acceleration cards to process large amounts of heterogeneous sensor data in real time and efficiently, ensuring low-latency system response. At the same time, it can also improve the computing performance of the fusion algorithm and enhance positioning stability and accuracy. At the same time, it can also be equipped with a dedicated motion sensor fusion chip (such as the VectorNav VN-100 series) for collaborative processing, which can effectively suppress IMU drift errors and improve the stability of posture estimation in complex terrain. At the same time, it can ensure that the robot arm's movements are based on real-time high-precision spatial posture, reducing picking errors. Based on the fused visual and inertial data, the robot's six-degree-of-freedom posture (position, velocity, and posture quaternion) can be predicted in real time, providing accurate spatial positioning information for robot arm control and path planning.
[0075] The defect detection module is electrically connected to the robotic arm control module and is used to detect and perform primary classification of internal defects of pineapple fruits based on multispectral-structured light data, and send the classification results to the robotic arm control module.
[0076] The defect detection module can be implemented based on the CPU / GPU resources in the posture prediction module. It can use the pre-trained multimodal deep learning model to analyze the multispectral and structured light fusion features, realize non-destructive detection and primary classification of internal defects in pineapple fruits, and identify defects such as hollowness, rot, and insect bites, thereby achieving high-accuracy internal quality inspection of the fruit, improving the quality control of picked fruits, and reducing the cost of manual post-sorting.
[0077] The robotic arm control module is used to dynamically adjust the picking strategy and the robot's robotic arm movements based on the posture prediction results and classification results.
[0078] The robotic arm control module can be constructed using an ABB IRC5 series industrial robotic arm controller, a CX2040 series motion controller, and an NI myRIO series real-time control system. It dynamically plans the robotic arm's motion trajectory and picking sequence based on posture prediction and defect detection classification results, allowing for efficient and accurate fruit picking operations. This improves the accuracy and flexibility of the robotic arm's motion, reduces the damage rate of the picking machinery, and improves overall picking efficiency and automation by dynamically adjusting the picking strategy.
[0079] The path planning module is electrically connected to the drive control module and is used to generate the robot's motion path in complex terrain based on the robot's positioning data and preset environmental three-dimensional data and send it to the drive control module.
[0080] The path planning module can also be implemented based on the CPU / GPU resources in the posture prediction module. At the same time, it is equipped with a path planning software framework that supports GPU acceleration (such as ROS Navigation Stack and MoveIt!). It can use a heuristic search algorithm to plan obstacle avoidance paths based on the robot's real-time positioning and pre-stored three-dimensional terrain model of the environment, ensuring the robot's safe and efficient movement in complex terrain, thereby achieving dynamic obstacle avoidance and adaptive terrain walking, improving operational safety, and optimizing motion paths, shortening picking time, and reducing energy consumption.
[0081] The drive control module is used to control the robot to move according to the received motion path.
[0082] The drive control module can adopt the Maxon EPOS4 series wheel / track drive controller, combined with a low-latency motor driver and encoder feedback system. It can drive the robot chassis to perform precise movements according to the instructions of the path planning module, achieve path tracking and area coverage, and realize smooth and precise robot motion control to ensure that the robot efficiently completes the coverage of the picking area and the continuity of operations.
[0083] The posture prediction module uses a combination of extended Kalman filtering and nonlinear optimization methods to estimate the robot posture when fusing visual data and inertial measurement data. The specific logic is:
[0084] Temporal synchronization and spatial calibration of the visual data and inertial measurement data collected by the robot;
[0085] The robot's posture is predicted based on the extended Kalman filter and updated in combination with visual data;
[0086] Based on nonlinear optimization method, the robot's historical posture is globally optimized to reduce the cumulative error;
[0087] The final output is the real-time position and posture prediction of the robot in complex terrain.
[0088] The robot's posture prediction is defined as the following vector form:
[0089]
[0090] In the formula represents the robot's posture vector at time t, p(t), v(t), q(t) represent the robot's position, velocity, and posture quaternion at time t, respectively. a (t), b g(t) represents the bias of acceleration and angular velocity, respectively. These two biases can be obtained by the accelerometer and gyroscope in the inertial measurement unit. (t) represents the accelerometer and gyroscope, respectively. (t) represents the system error components existing in measuring acceleration and angular velocity, respectively. They are manifested as constant or slowly changing offsets of the sensor output. These two biases are key parameters in the error model of the inertial measurement unit (i.e., IMU). Real-time estimation and correction through state estimation algorithms (such as extended Kalman filtering) can reduce the impact of sensor errors on the robot's spatial positioning and posture estimation, thereby ensuring positioning accuracy and system stability.
[0091] Among them, attitude quaternion is a mathematical tool used to represent the rotation state of objects (such as robots and robotic arms) in three-dimensional space. It is used to describe the rotation relationship from the reference coordinate system (usually the ground or world coordinate system) to the object's own coordinate system. Compared with traditional Euler angles or rotation matrices, quaternions are more efficient in calculation and avoid "universal lock".
[0092] (Gimbal Lock) and other problems, is a posture representation method widely used in modern robotics and aerospace fields.
[0093] Specifically, a posture quaternion can be expressed as:
[0094]
[0095] The q here w represents the scalar part of the real part, q x ,q y ,q z The vector part represents the imaginary part. i, j, and k are different from the subscripts in the following text. Here they represent the imaginary unit of the imaginary part.
[0096] Assume that the rotation axis U=[u x ,u y ,u z ] T When rotating an angle ∈, the corresponding attitude quaternion can be expressed as:
[0097]
[0098] Attitude quaternions can represent the robot's current spatial attitude (i.e., rotation state), ensuring that the robot can accurately perceive its own direction in complex terrain, thereby providing reliable attitude information for robotic arm control and path planning.
[0099] When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing the cost function, which is expressed as follows:
[0100]
[0101] Where r(IMU,t) represents the IMU pre-integration residual at time t, r(vision,t) represents the visual measurement residual at time t, ∑IMU,t and ∑vision,t represent the covariance matrices of the IMU pre-integration residual and the visual measurement residual at time t, respectively. represents the weighted Euclidean norm.
[0102] From the expression of the cost function, it can be seen that the cost function aims to estimate the robot state vector so that the residual (error) of combining the IMU pre-integrated data and the visual measurement data is minimized to achieve optimal spatial positioning and posture estimation.
[0103] Specifically, IMU data can provide continuous instantaneous acceleration and angular velocity information, and infer motion trajectories through inertial navigation, but its integration process will accumulate errors, especially in the presence of noise and bias; while visual data provides relative pose measurement (such as visual odometry), and achieves pose constraints through image feature matching, which can effectively suppress IMU error accumulation. However, it is affected by occlusion and lighting. The cost function weighs the residuals of these two types of data so that the positioning results take into account both the dynamic continuity of the IMU and the geometric constraints of visual measurements, thereby improving robustness and accuracy.
[0104] The IMU uses accelerometer and gyroscope measurements to infer the robot's predicted state. The difference between this and the current optimized variable state is the residual. The IMU residual reflects the degree to which the state variables explain the inertial measurements. State estimates that do not conform to the IMU motion model will result in large residuals. Visual residuals, derived from the relative pose estimated by matching feature points between consecutive image frames, differ from the pose at the corresponding time point in the optimized state. They constrain the robot's position and attitude and help correct drift in the IMU integration. The covariance matrix of the two is a statistical description of measurement noise. Higher-precision sensors have smaller covariances, and these residuals are given greater weight during optimization, achieving reasonable data fusion. Therefore, by minimizing the weighted sum of squared residuals, the algorithm leverages the spatial constraints of visual measurements while taking into account the dynamic continuity of the IMU. This yields a robot state estimate that best matches multi-source observations, ensuring the robot can obtain highly accurate and robust position and attitude information in complex terrain environments, thereby supporting subsequent precise picking and path planning by the robotic arm.
[0105] The defect detection module has a built-in multimodal deep learning model for detecting internal defects in pineapples. The training logic of the multimodal deep learning model is as follows:
[0106] Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the categories of the internal defects were annotated.
[0107] Construct a fusion feature vector of multispectral-structured light and input it into a multimodal deep learning model for training;
[0108] Using cross entropy as the loss function, with the goal of minimizing the loss function, the model parameters of the multimodal deep learning model are repeatedly iterated until optimization is complete. During iteration, forward propagation (calculating output) and backpropagation (calculating gradients and updating parameters) are repeated until the loss function converges or the preset number of iterations is reached.
[0109] The purpose of this step of the training process is to establish a labeled dataset covering various defect types (such as hollowness, rot, insect damage, etc.) to ensure the diversity and representativeness of the training samples.
[0110] in:
[0111] The expression of the fused feature vector is:
[0112]
[0113] In the formula Represents the fusion feature vector of the pixel at the coordinate (x, y), I(λ i ,x,y) indicates the wavelength is λ i Where D(x,y) represents the multispectral reflectance of the pixel at (x,y), D(x,y) represents the structured light depth value of the pixel at (x,y), k(x,y) represents the curvature characteristic of the pixel at (x,y), x and y represent the horizontal and vertical coordinates of the pixel in the image, the subscript i represents the index of the selected wavelength, and N is the total number of selected wavelengths. When selecting wavelengths, one should choose a band that maximizes the difference in reflectance between internal defects (such as decay and lesions) and healthy tissue to improve the sensitivity and accuracy of defect detection. For example, moisture (700nm-900nm), chlorophyll (670nm-680nm), and carotenoids (560nm-590nm) exhibit significant differences between specific near-infrared and visible light bands, and these wavelengths should be preferred.
[0114] From the expression of the fusion feature vector, it can be seen that it is divided into three parts. The first part is the multispectral reflectance, which is used to reflect the differences in the internal components of the fruit (such as water, sugar, pigment, etc.) and the surface state; the second part is the structured light depth feature, which is used to reflect the three-dimensional spatial position and morphological information of the pixel point corresponding to the fruit surface, so as to reveal the spatial characteristics such as depression and expansion inside the fruit; the third part is the curvature feature, which can be obtained through the structured light depth map, which is used to describe the local curvature of the fruit surface, that is, the surface shape change rate, so as to locate the defect boundary and identify abnormalities on the fruit surface.
[0115] When weighting the fused feature vector, the self-attention mechanism is introduced to extract the contextual latent state features of each pixel point, and the dynamic weight of each vector element is calculated based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as:
[0116]
[0117] Where w ∈ (x,y) represents the dynamic weight of the ∈th vector element in the fusion feature vector, h(x,y) represents the contextual hidden state feature of the pixel at coordinate (x,y), ∈ represents the index of the vector element, and f qz (·) represents a weight generation function. It can be understood that since the vector elements in the fusion feature vector include N groups of multispectral reflectances, a group of structured light depth values, and a group of curvature features, the total number of elements is N+2.
[0118] The contextual latent state feature represents the contextual information state of the current pixel. It can capture the surrounding environment and semantic information of the pixel, and help the weight generation function understand the importance of the location feature. Specifically, the semantic and spatial information of the pixel and its neighborhood can be extracted through the intermediate layer feature map of the multimodal deep learning model, and then combined with the global context information generated by the self-attention mechanism (Transformer, non-local network, etc.), and finally obtained through feature fusion. It can facilitate the identification of environmental factors such as lighting changes and occlusions, adjust the modal weights to reduce the contribution of the affected modal, and combine the neighborhood defect distribution information to enhance the smoothness and consistency of the modal weights in the local continuous area, thereby dynamically adjusting the sensitivity of different modalities to specific defect types and achieving more accurate defect positioning.
[0119] The weight generation function can be designed as a neural network module with parameters. Specifically, its expression can be set as:
[0120]
[0121] Where A and B represent the weight and bias of the hidden layer respectively, a ∈ 、b ∈ They represent the weight and bias of the output layer corresponding to the ∈th vector element, σ(·) represents the activation function, which can be ReLu or LeakyReLu, etc. This structure allows dynamic weights to capture the nonlinear relationship between context and features, dynamically reflecting the importance of the environmental information at that location and the current feature.
[0122] Then use the dynamic weight to perform weighted processing on the fusion feature vector, and the calculation method is:
[0123]
[0124] In the formula represents the fused feature vector after weighted processing, ⊙ represents the element-by-element product;
[0125] The expression of the cross entropy loss function is:
[0126]
[0127] Where θ represents the model parameters to be trained, y m 、 They represent the true label and model prediction label of the internal defect category of the sample, the subscript m represents the index of the sample, and λ represents the regularization hyperparameter. Representing the spatial gradient of weights can smooth the weight distribution and improve the robustness of the model.
[0128] Cross entropy measures the difference between the model's predicted probability distribution and the true label distribution. By continuously adjusting the parameter θ through optimization algorithms such as gradient descent, the predicted probability gradually approaches the true label, thereby improving the classification accuracy.
[0129] Specifically, the multimodal deep learning model is a holistic machine learning framework or system designed to process and fuse multi-source heterogeneous data (such as multispectral image data and structured light depth information) to extract comprehensive features and complete the target task (such as internal defect detection and classification). The classifier can use a neural network discriminant function to classify or discriminate the fused features and output the defect category and the corresponding class probability.
[0130] The logic for primary classification of pineapple fruits using the trained multimodal deep learning model is:
[0131] The multispectral-structured light data of pineapple fruits collected in real time are fused to obtain the fusion feature vector of multispectral-structured light.
[0132] Input the fused feature vector into the multimodal deep learning model and output the corresponding internal defect category and defect probability;
[0133] The product of the severity of each internal defect category and the probability of the defect is used as a defect score, and the quality of the pineapple fruit is graded based on the defect score. The severity of each internal defect can be determined based on expert experience, with larger values indicating more severe defects. For example, the severity of hollowness is 1, the severity of insect damage is 2, and the severity of rot is 3. A higher defect score indicates a lower quality grade, while a lower value indicates a higher quality grade.
[0134] In this step, by combining multispectral optical information and structured light geometric information, the shortcomings of a single modality are supplemented, the classification accuracy is significantly improved, and the detection accuracy and robustness are higher. At the same time, the defect score is calculated by probability weighted severity to avoid the binary judgment of simple hard classification, reflecting the severity and confidence of the defect. Multiple categories of defects and corresponding weights can be flexibly set to meet the quality management needs of different fruits.
[0135] The logic of the robotic arm control module to dynamically adjust the picking strategy and the robot's robotic arm movements based on the posture prediction results and classification results is as follows:
[0136] Determine the position and posture of the robotic arm base based on the real-time position and posture prediction of the robot;
[0137] The robot arm's picking path is planned based on the quality grade of the pineapple fruit, with priority given to picking pineapples with higher quality grades.
[0138] After each picking is completed, the picking order is dynamically adjusted and the picking path of the robotic arm is updated until all pineapple fruits in the area are picked.
[0139] It can be understood that the manipulator is installed on a fixed rigid structure of the robot (usually the robot chassis, body or frame), so its posture has a fixed rigid transformation relationship with the posture of the robot. Suppose the fixed rigid body transformation of the manipulator base relative to the robot body coordinate system is:
[0140]
[0141] Where T b / r represents the transformation matrix, R b / r represents the rotation matrix of the manipulator base relative to the robot coordinate system, p b / r Represents the translation vector of the manipulator base relative to the robot coordinate system, R b / r 、p b / r are all fixed constants;
[0142] Then convert the robot position and posture quaternion into a homogeneous transformation matrix:
[0143]
[0144] Where T r (t) represents the homogeneous transformation matrix, R r (t) represents the rotation matrix obtained by transforming the attitude quaternion q(t);
[0145] According to the rigid body transformation chain, the posture vector of the robot base in the environment coordinate system can be expressed as:
[0146] T b (t) = T r(t)·T b / r
[0147] Right now:
[0148]
[0149] If:
[0150] R b (t) = R r (t)·R b / r
[0151] p b (t) = R r (t)·p b / r +p(t)
[0152] Then the posture vector of the robot arm base can be expressed as:
[0153]
[0154] Here R b (t) is a rotation matrix, which represents the rotational posture of the manipulator base relative to the global environment coordinate system at time t, that is, the orientation and posture direction of the manipulator base; p b (t) is a three-dimensional vector, which represents the position coordinates of the robot base in the global environment coordinate system at time t, that is, the position of the spatial point where the robot base is located.
[0155] It can be seen that the attitude vector of the manipulator base is the combination of the position and attitude in the robot attitude vector and the fixed transformation of the base relative to the robot. In other words, the global posture of the robot determines the spatial position of the manipulator base, and the fixed geometric parameters of the manipulator base relative to the robot body determine the rigid body transformation relationship between the two.
[0156] Then the location information and quality grade information of pineapple fruits are divided into a set S, which can be expressed as:
[0157] S={(Pg j ,L j )|j=1,2,…,M}
[0158] Where Pg j Indicates the position of the j-th pineapple fruit, L j represents the defect score of the j-th pineapple fruit, M represents the total number of pineapple fruits to be picked in the area, and j represents the index of the pineapple fruit.
[0159] Score L for the set S according to the defect score j Arrange them in ascending order, and the sorted set is represented as S′, in which the pineapple fruit grades are arranged from high to low, so that fruits with better quality can be picked first.
[0160] For the sorted set S, the fruits are picked in order, and the picking path of the robot arm can be expressed as a sequence Q:
[0161] Q=[Pg j1 ,Pg j2 ,…,Pg jq ,…,Pg jm ]
[0162] Where Pg jq It represents the position of the qth picking target, and the subscript jq represents the index of the picking target, which is obtained from the sorted set S′.
[0163] When updating and optimizing the picking path, the goal is to minimize the total length of the path, so the optimization function can be expressed as:
[0164]
[0165] This formula minimizes the sum of the distances between two adjacent picking points in the path, aiming to reduce the total distance the robotic arm moves and improve efficiency.
[0166] In this step, by combining the robot's global posture vector with the rigid body transformation of the robotic arm base, the real-time position of the robotic arm base in the environment is accurately obtained, ensuring that the robotic arm movement is based on accurate spatial reference, reducing motion errors and collision risks, and then using the fruit defect score to intelligently sort the picking targets. Compared with the traditional robotic arm positioning method for picking, this method of using fruit quality as a reference to correct the robotic arm positioning and motion trajectory can achieve "prioritized picking of high-quality fruits", thereby improving the overall quality of the output fruit, reducing the picking of low-quality fruits, and reducing resource waste.
[0167] The planning logic of the robot's motion path is:
[0168] Identify the robot's traversable area during movement based on the robot's position data and a preset terrain model;
[0169] Plan the shortest path from the current area to the next area based on a heuristic search algorithm, avoiding obstacles and steep slopes;
[0170] Generate continuous motion trajectories until all pineapple fruits in the area are picked.
[0171] Combining the robot's real-time position data with a pre-built three-dimensional terrain model can accurately identify the robot's traversable areas in complex terrain, avoid untravable paths caused by blind planning, and improve the safety and practicality of path planning.
[0172] In this example, the multimodal solution of the present invention is compared with two single-modal solutions (single IMU solution and single vision solution) under traditional methods. Pineapple fruits in 20 areas are picked respectively. After each area is picked, the position error, posture error, defect accuracy rate, picking path and other parameters are calculated. The parameters of the overall picking process are as follows:
[0173] Table 1: Multimodal scheme picking parameters
[0174]
[0175]
[0176] Table 2: Single modality (IMU) solution selection parameters
[0177]
[0178]
[0179] Table 3: Picking parameters for single-modality (vision) solution
[0180]
[0181]
[0182] It is understandable that the single IMU solution cannot detect defects in fruits, so manual inspection is performed and the defect detection accuracy is calculated. The single vision solution uses a traditional visual recognition algorithm (YoLoV5 model) to detect whether there are defects. Since the single multispectral-structured light solution only involves fruit defect recognition and does not involve kinematic directions such as position estimation and posture estimation, no comparison of the solutions is set.
[0183] Please refer to Figure 3 , after combining the parameters in the above table, the maximum and minimum normalization method is performed to eliminate the dimension. Specifically, for the parameters where the smaller the better (i.e., position estimation error, attitude estimation error, and total length of the picking path), the normalized calculation formula is:
[0184]
[0185] For the parameter where the larger the better, the normalized calculation formula is:
[0186]
[0187] After normalization, the data in the 20 regions can be averaged and plotted.
[0188] The advantage of doing this is that it not only eliminates dimensions but also eliminates differences in decision logic. That is, the closer the normalized values of all parameters are to 1, the better the performance, and the closer they are to 0, the worse the performance. In this way, a bar chart can be constructed to intuitively show the differences between the three different solutions.
[0189] From the table data and bar chart, we can see that although IMU (Inertial Measurement Unit) can continuously output acceleration and angular velocity information, the error will accumulate rapidly over time when its data is integrated and calculated for positioning (drift phenomenon). Especially in complex terrain and long-term operation scenarios, the positioning error is often large. It relies on IMU data and lacks auxiliary correction of external sensors such as vision and GPS. It cannot directly perceive the environment. The lack of visual information leads to weak path planning and obstacle avoidance capabilities, which easily leads to unreasonable path planning and obstacle avoidance failure. At the same time, it is also impossible to detect fruit defects. Therefore, the performance of its various parameters is very poor. The performance of the single-vision solution is low, and its applicability in actual application scenarios is very poor; the performance of the single-vision solution is better in comparison, because the visual camera can obtain rich environmental information, including color images and depth information (RGB-D), and support more accurate feature extraction and environmental modeling, but the positioning stability and real-time performance are insufficient in dynamic or complex terrain, affecting the continuity of robot arm control and path planning; the multimodal solution of the present invention has better performance in various parameters. By fusing multiple heterogeneous data, it can dynamically adjust the robot arm's motion trajectory and picking strategy, achieve priority picking of high-quality fruits, reduce the risk of mechanical damage, and improve production efficiency.
[0190] A spatially heterogeneous visual positioning method for a pineapple picking robot is provided. The positioning method is applicable to the above-mentioned positioning system and specifically comprises the following steps:
[0191] S1: Collects multimodal visual data and inertial measurement data of the complex terrain environment in which the robot is located, and simultaneously collects multispectral and structured light data of the target fruit;
[0192] S2: The robot achieves spatial positioning and posture prediction through a visual-inertial fusion algorithm. It also uses multispectral and structured light data to detect and classify internal fruit defects in real time based on a multimodal deep learning model.
[0193] S3: Control the picking path of the robotic arm based on the robot's position, posture, and pineapple fruit quality level to complete the picking of pineapples in the area where the robot is located;
[0194] S4: The robot's motion path is planned according to the robot's position data and a preset terrain model, and steps S1 to S3 are repeated in each area until the pineapple fruits in all areas are picked.
[0195] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0196] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software depends on the specific application and design constraints of the technical solution.
[0197] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.
[0198] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A spatially heterogeneous visual positioning system for a pineapple picking robot, characterized in that: Specifically include: a data acquisition module, electrically connected to the posture prediction module and the path planning module, for collecting visual data, inertial measurement data, position data, and multispectral-structured light data of the robot in complex terrain, and sending the data to the posture prediction module; A posture prediction module, electrically connected to the manipulator control module, configured to predict the robot's posture based on the visual data and the inertial measurement data, and to send the posture prediction result to the manipulator control module; a defect detection module, the defect detection module being electrically connected to the robotic arm control module and configured to detect and perform preliminary classification of internal defects of the pineapple fruit based on the multispectral-structured light data, and to send the classification results to the robotic arm control module; A robotic arm control module, which is used to dynamically adjust the picking strategy and the robot's robotic arm movements based on the posture prediction results and the classification results; A path planning module, electrically connected to the drive control module, for generating a motion path for the robot in complex terrain based on the robot's positioning data and preset three-dimensional environmental data, and sending the generated path to the drive control module; A drive control module is used to control the robot to move according to the received motion path.
2. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 1, characterized in that: The visual data is an RGB-D image of the robot's location; the inertial measurement data includes the robot's speed, acceleration, and angular velocity during movement; the multispectral-structured light data includes multispectral images and structured light depth maps, and covers near-infrared and short-wave infrared bands.
3. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 2, characterized in that: The posture prediction module uses a combination of extended Kalman filtering and nonlinear optimization methods to estimate the robot posture when fusing visual data and inertial measurement data. The specific logic is: Temporal synchronization and spatial calibration of the visual data and inertial measurement data collected by the robot; The robot's posture is predicted based on the extended Kalman filter and updated in combination with visual data; Based on nonlinear optimization method, the robot's historical posture is globally optimized to reduce the cumulative error; The final output is the real-time position and posture prediction of the robot in complex terrain.
4. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 3, wherein: The robot's posture prediction is defined as the following vector form: In the formula represents the robot's posture vector at time t, p(t), v(t), q(t) represent the robot's position, velocity, and posture quaternion at time t, respectively. a (t), b g (t) represents the bias of acceleration and angular velocity respectively; When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing the cost function, which is expressed as follows: Where r(IMU,t) represents the IMU pre-integration residual at time t, r(vision,t) represents the visual measurement residual at time t, ∑IMU,t and ∑vision,t represent the covariance matrices of the IMU pre-integration residual and the visual measurement residual at time t, respectively. represents the weighted Euclidean norm.
5. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 4, characterized in that: The defect detection module has a built-in multimodal deep learning model for detecting internal defects of pineapple fruits. The training logic of the multimodal deep learning model is: Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the categories of the internal defects were annotated. Construct a fusion feature vector of multispectral-structured light, perform weighted processing on it, and then input it into a multimodal deep learning model for training; Using cross entropy as the loss function, with the goal of minimizing the loss function, the model parameters of the multimodal deep learning model are repeatedly iterated until the optimization is completed; in: The expression of the fused feature vector is: In the formula Represents the fusion feature vector of the pixel at the coordinate (x, y), I(λ i ,x,y) indicates the wavelength is λ i The multispectral reflectance of the pixel at coordinate (x, y) is represented by D(x, y), the structured light depth value of the pixel at coordinate (x, y), k(x, y) represents the curvature feature of the pixel at coordinate (x, y), x and y represent the horizontal and vertical coordinates of the pixel in the image, the subscript i represents the index of the selected wavelength, and N is the total number of selected wavelengths; When weighting the fused feature vector, the self-attention mechanism is introduced to extract the contextual latent state features of each pixel point, and the dynamic weight of each vector element is calculated based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as: Where w ∈ (x,y) represents the dynamic weight of the ∈th vector element in the fusion feature vector, h(x,y) represents the contextual hidden state feature of the pixel at coordinate (x,y), ∈ represents the index of the vector element, and f qz (·) represents the weight generating function; Then use the dynamic weight to perform weighted processing on the fusion feature vector, and the calculation method is: In the formula represents the fused feature vector after weighted processing, ⊙ represents the element-by-element product; The expression of the cross entropy loss function is: Where θ represents the model parameters to be trained, y m 、 They represent the true label and model prediction label of the internal defect category of the sample, the subscript m represents the index of the sample, and λ represents the regularization hyperparameter. represents the spatial gradient of the weights.
6. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 5, characterized in that: The logic for primary classification of pineapple fruits using the trained multimodal deep learning model is: The multispectral-structured light data of pineapple fruits collected in real time are fused to obtain the fusion feature vector of multispectral-structured light. The fused feature vector is weighted and input into the multimodal deep learning model to output the corresponding internal defect category and defect probability; The product of the severity of the internal defect category and the defect probability is used as the defect score, and the quality of pineapple fruit is graded according to the defect score.
7. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 6, characterized in that: The logic of the robotic arm control module dynamically adjusting the picking strategy and the robot's robotic arm motion according to the posture prediction results and classification results is as follows: Determine the position and posture of the robotic arm base based on the real-time position and posture prediction of the robot; The robot arm's picking path is planned based on the quality grade of the pineapple fruit, with priority given to picking pineapples with higher quality grades. After each picking is completed, the picking order is dynamically adjusted and the picking path of the robotic arm is updated until all pineapple fruits in the area are picked.
8. The spatially heterogeneous visual positioning system for a pineapple picking robot according to claim 7, characterized in that: The planning logic of the robot motion path is: Identify the robot's traversable area during movement based on the robot's position data and a preset terrain model; Plan the shortest path from the current area to the next area based on a heuristic search algorithm, avoiding obstacles and steep slopes; Generate continuous motion trajectories until all pineapple fruits in the area are picked.
9. A spatially heterogeneous visual positioning method for a pineapple picking robot, characterized by: The positioning method is applicable to the positioning system according to any one of claims 1 to 8, and the specific steps include: S1: Collects multimodal visual data and inertial measurement data of the complex terrain environment in which the robot is located, and simultaneously collects multispectral and structured light data of the target fruit; S2: The robot achieves spatial positioning and posture prediction through a visual-inertial fusion algorithm. It also uses multispectral and structured light data to detect and classify internal fruit defects in real time based on a multimodal deep learning model. S3: Control the picking path of the robotic arm based on the robot's position, posture, and pineapple fruit quality level to complete the picking of pineapples in the area where the robot is located; S4: The robot's motion path is planned according to the robot's position data and a preset terrain model, and steps S1 to S3 are repeated in each area until the pineapple fruits in all areas are picked.
Citation Information
Patent Citations
Litchi recognition method based on visual algorithm and bionic litchi picking robot
CN115553132A
Cooperative operation method, system and platform for guiding robot to pick and convey fruits
CN117546681A
Spatial heterogeneous visual positioning system of tomato picking robot
CN119772891A
Intelligent picking robot and monitoring system
CN119949151A
Cotton picking robot and picking navigation method thereof
CN120167228A
Cited By
Robot hybrid visual servo positioning method, device and equipment
CN121004614A
Flying robot system for areca nut picking and control method thereof
CN121100681A
Flying robot system for areca picking and control method thereof
CN121100681B
Picking robot and fruit pose detection and interactive visualization method and system thereof
CN121391976A
Dynamic fruit detection tracking and real-time positioning method of citrus picking robot
CN122244107A