A spatial heterogeneous visual positioning system and positioning method for a pineapple picking robot
By using a spatial heterogeneous visual positioning system for pineapple harvesting robots that integrates multi-source information, the problems of positioning accuracy and fruit defect detection in complex terrain have been solved, enabling efficient and stable pineapple harvesting and improving harvesting quality and efficiency.
Patent Information
- Application Number
- CN202510860889.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing pineapple harvesting robots lack sufficient positioning accuracy and robustness in complex terrain, making it difficult to efficiently and non-destructively detect internal defects in the fruit, and resulting in low harvesting efficiency and quality.
A spatial heterogeneous visual positioning system for pineapple harvesting robots, employing multi-source information fusion, combines visual, inertial measurement, multispectral, and structured light data. It uses extended Kalman filtering and nonlinear optimization methods for attitude prediction, utilizes a multimodal deep learning model to detect fruit defects, and combines path planning algorithms to optimize harvesting strategies.
It achieves high-precision spatial positioning and attitude estimation in complex terrain, improves harvesting stability and real-time performance, enhances the accuracy of fruit internal defect detection and harvesting quality, optimizes harvesting paths, and reduces the risk of mechanical damage.
Smart Images

Figure CN120593732B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of positioning technology for harvesting robots, specifically to a spatial heterogeneous visual positioning system and positioning method for a pineapple harvesting robot. Background Technology
[0002] With the acceleration of agricultural modernization, intelligent harvesting robots are increasingly widely used in fruit and vegetable harvesting, especially in automated harvesting tasks under complex terrain conditions. The accuracy and robustness of the robot's positioning and control system have become key technological bottlenecks. Traditional harvesting robots mostly rely on single vision or inertial sensors for environmental perception and attitude estimation, which is insufficient to meet the high-precision spatial positioning requirements of uneven terrain such as terraces and slopes. This leads to increased positioning errors in the robotic arm, unstable harvesting movements, and affects harvesting efficiency and fruit quality. Furthermore, the robot lacks efficient and non-destructive detection capabilities for internal fruit defects, hindering intelligent grading and harvesting, and significantly falling short in improving harvesting quality and reducing subsequent processing costs.
[0003] The existing technology, disclosed in CN119772891A, discloses a spatial heterogeneous visual positioning system for a tomato harvesting robot, involving the field of deep learning. The system includes the following steps: S1, calibrating the positioning of the tomato harvesting robot; S2, acquiring tomato images in an open environment using a fixed camera; S3, acquiring the spatial coordinate information of the fruit under the fixed camera's field of view; S4, determining the reachable space of the harvesting robot arm and moving the harvesting robot arm; S5, using a palm-sized camera to perform secondary spatial information perception on the target, further acquiring spatial coordinate information in a relatively closed environment; S6, performing approximate sphere fitting to obtain the optimal harvesting target; and S7, determining the number of tomatoes and estimating the harvesting pose. While this system can improve positioning accuracy, it relies on multiple precise calibrations and multi-stage visual perception, resulting in low real-time performance and robustness. Furthermore, it lacks multi-sensor fusion, and its intelligent fusion and dynamic adaptation capabilities are insufficient, making it difficult to guarantee high-precision positioning and stable harvesting in complex environments, thus limiting operational efficiency and environmental adaptability.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a spatial heterogeneous visual positioning system and positioning method for a pineapple harvesting robot, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A spatial heterogeneous visual positioning system for a pineapple harvesting robot, specifically comprising:
[0008] The data acquisition module is electrically connected to the attitude prediction module and the path planning module. It is used to collect the robot's visual data, inertial measurement data, position data and multispectral-structured light data of pineapple fruit in complex terrain, and send them to the attitude prediction module.
[0009] The attitude prediction module is electrically connected to the robotic arm control module. It is used to predict the robot's attitude based on visual data and inertial measurement data, and send the attitude prediction results to the robotic arm control module.
[0010] The defect detection module is electrically connected to the robotic arm control module. It is used to detect and perform preliminary classification of internal defects of pineapple fruits based on multispectral structured light data, and send the classification results to the robotic arm control module.
[0011] A robotic arm control module is used to dynamically adjust the picking strategy and the robotic arm movements of the robot based on the posture prediction results and classification results.
[0012] The path planning module is electrically connected to the drive control module and is used to generate the robot's motion path in complex terrain based on the robot's positioning data and preset environmental 3D data, and send it to the drive control module.
[0013] A drive control module is used to control the robot to move according to the received motion path.
[0014] Preferably, the visual data is an RGB-D image of the robot's location; the inertial measurement data includes the robot's velocity, acceleration, and angular velocity during movement; and the multispectral-structured light data includes multispectral images and structured light depth maps, covering near-infrared and short-wave infrared bands.
[0015] Preferably, when fusing visual data and inertial measurement data, the posture prediction module uses a combination of extended Kalman filtering and nonlinear optimization methods to estimate the robot's posture. The specific logic is as follows:
[0016] Time synchronization and spatial calibration are performed on the visual data and inertial measurement data collected by the robot;
[0017] Robot pose prediction is performed based on extended Kalman filtering, and pose update is performed by combining visual data.
[0018] The robot's historical poses are globally optimized using a nonlinear optimization method, thereby reducing accumulated errors.
[0019] The final output is a prediction of the robot's real-time position and attitude in complex terrain.
[0020] Preferably, the robot's pose prediction is defined in the following vector form:
[0021]
[0022] In the formula Let p(t), v(t), and q(t) represent the robot's pose vector at time t, respectively, and let b be the robot's position, velocity, and pose quaternions at time t. a (t), b g (t) represent the biases for acceleration and angular velocity, respectively;
[0023] When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing a cost function, the expression of which is shown below:
[0024]
[0025] In the formula, r(IMU,t) represents the IMU pre-integration residual at time t, r(vision,t) represents the visual measurement residual at time t, and ∑IMU,t and ∑vision,t represent the covariance matrices of the IMU pre-integration residual and the visual measurement residual at time t, respectively. This represents the weighted Euclidean norm.
[0026] Preferably, the defect detection module incorporates a multimodal deep learning model for detecting internal defects in pineapple fruits. The training logic of the multimodal deep learning model is as follows:
[0027] Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the types of internal defects were labeled.
[0028] A multispectral-structured light fusion feature vector is constructed, weighted, and then input into a multimodal deep learning model for training.
[0029] Cross-entropy is used as the loss function, and the model parameters of the multimodal deep learning model are iterated repeatedly until the optimization is completed, with the goal of minimizing the loss function.
[0030] in:
[0031] The expression for fusing feature vectors is:
[0032]
[0033] In the formula I(λ) represents the fused feature vector of the pixel at coordinates (x, y). i (x, y) represents wavelength λ i The time coordinate is the multispectral reflectance of the pixel at (x,y), D(x,y) represents the structured light depth value of the pixel at (x,y), k(x,y) represents the curvature feature of the pixel at (x,y), x and y represent the horizontal and vertical coordinates of the pixel in the image, the subscript i represents the index of the selected wavelength, and N is the total number of selected wavelengths.
[0034] When weighting the fused feature vector, a self-attention mechanism is introduced to extract the contextual latent state features of each pixel, and to calculate the dynamic weight of each vector element based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as follows:
[0035]
[0036] In the formula w ∈ (x,y) represents the dynamic weight of the ∈-th vector element in the fused feature vector, h(x,y) represents the contextual latent state feature of the pixel at coordinates (x,y), ∈ represents the index of the vector element, and f qz (·) represents the weight generation function;
[0037] The fused feature vector is then weighted using dynamic weights, calculated as follows:
[0038]
[0039] In the formula represents the weighted fused feature vector, and ⊙ represents the element-wise product;
[0040] The expression for the cross-entropy loss function is:
[0041]
[0042] In the formula, θ represents the model parameters to be trained, and y m , λ and λ represent the true label and model predicted label of the internal defect category of the sample, respectively. The subscript m represents the sample index, and λ represents the regularization hyperparameter. Represents the spatial gradient of the weights.
[0043] Preferably, the logic for using a trained multimodal deep learning model to perform primary classification of pineapple fruits is as follows:
[0044] Feature fusion was performed on the multispectral-structured light data of pineapple fruits collected in real time to obtain the fused feature vector of multispectral-structured light.
[0045] The fused feature vectors are weighted and then input into a multimodal deep learning model, which outputs the corresponding internal defect category and defect probability.
[0046] The product of the severity of the internal defect category and the defect probability is used as the defect score, and the quality of the pineapple fruit is graded based on the defect score.
[0047] Preferably, the logic of the robotic arm control module dynamically adjusting the picking strategy and the robot's robotic arm movements based on the posture prediction results and classification results is as follows:
[0048] The position and orientation of the robotic arm base are predicted based on the robot's real-time position and orientation.
[0049] The robotic arm's harvesting path is planned according to the quality grade of the pineapple fruit, with priority given to harvesting pineapple fruits of higher quality grade.
[0050] After each harvest, the harvesting order is dynamically adjusted and the robotic arm's harvesting path is updated until all pineapple fruits in the area have been harvested.
[0051] Preferably, the planning logic for the robot's motion path is as follows:
[0052] Identify the robot's traversable areas during its movement based on the robot's location data and a pre-set terrain model;
[0053] The algorithm uses a heuristic search algorithm to plan the shortest path from the current location to the next location, while avoiding obstacles and steep slopes.
[0054] Generate a continuous motion trajectory until all pineapple fruits in all areas have been harvested.
[0055] A spatial heterogeneous visual positioning method for a pineapple harvesting robot, wherein the positioning method is applicable to the aforementioned positioning system, and the specific steps include:
[0056] S1: Collect multimodal visual data and inertial measurement data of the robot in the complex terrain environment, and at the same time collect multispectral and structured light data of the target fruit;
[0057] S2: The robot achieves spatial localization and attitude prediction through a vision-inertial fusion algorithm. At the same time, it uses multispectral and structured light data to detect internal defects in fruits in real time and perform primary classification based on a multimodal deep learning model.
[0058] S3: Control the picking path of the robotic arm based on the robot's position, posture and the quality grade of the pineapple fruit to complete the pineapple picking in the area where the robot is located;
[0059] S4: Based on the robot's location data and the preset terrain model, plan the robot's movement path and repeat steps S1 to S3 in each area until all pineapple fruits in all areas have been harvested.
[0060] Compared with the prior art, the beneficial effects of the present invention are:
[0061] This invention achieves high-precision spatial positioning and attitude estimation for robots in complex terrain environments by comprehensively utilizing multi-source information such as vision, inertial measurement, multispectral, and structured light, significantly improving positioning stability and real-time performance. Through an embedded multimodal deep learning model, it enables efficient and non-destructive detection and intelligent classification of internal defects in pineapple fruits, supporting the robotic arm to dynamically adjust its harvesting strategy, prioritizing the harvesting of higher-quality fruits and improving harvesting quality and efficiency. Simultaneously, by combining a pre-set 3D terrain model with a heuristic search algorithm, it effectively plans the robot's motion path, ensuring safe and stable movement in complex environments and reducing the risk of mechanical damage. Attached Figure Description
[0062] Figure 1 This is a schematic diagram of the overall modular structure of the present invention;
[0063] Figure 2 This is a schematic diagram of the overall process of the present invention;
[0064] Figure 3 This is a parameter comparison chart of different schemes of the present invention. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0066] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0067] Example:
[0068] Please see Figures 1-3 The present invention provides a technical solution:
[0069] A spatial heterogeneous visual positioning system for a pineapple harvesting robot, specifically comprising: a data acquisition module, a posture prediction module, a defect detection module, a robotic arm control module, a path planning module, and a drive control module.
[0070] The data acquisition module is electrically connected to the attitude prediction module and path planning module. It is used to collect visual data, inertial measurement data, position data, and multispectral-structured light data of the pineapple fruit in complex terrain, and send them to the attitude prediction module. The visual data is an RGB-D image of the robot's location; the inertial measurement data includes the robot's velocity, acceleration, and angular velocity during movement; the multispectral-structured light data includes multispectral images and structured light depth maps, covering the near-infrared and short-wave infrared bands.
[0071] The data acquisition module can be composed of various hardware structures, such as: using a RealSense D455 series RGB-D camera to detect visual data, using a BMI 160 series inertial measurement unit to detect inertial measurement data, using a NEO-M8N series GPS positioning module to acquire the robot's position data, and using a RedEdge-MX series multispectral camera and a RealSense L515 series structured light sensor to jointly detect the multispectral-structured light data of the pineapple fruit.
[0072] The data acquisition module is responsible for acquiring multimodal data of the robot's operating environment, including color and depth images, robot motion inertial information, real-time geographical location, and multispectral and structured light information of pineapple fruits. The combination of multi-source data improves the accuracy and robustness of environmental perception, providing accurate raw data for subsequent data processing and decision-making.
[0073] The attitude prediction module is electrically connected to the robotic arm control module. It is used to predict the robot's attitude based on visual data and inertial measurement data, and then send the attitude prediction results to the robotic arm control module.
[0074] The attitude prediction module can utilize the AGX Xavier series of industrial-grade embedded computing platforms and can also be configured with Xilinx Zynq Ultrascale series FPGA acceleration cards to process large amounts of heterogeneous sensor data in real time and efficiently, ensuring low-latency system response. This also improves the computational performance of the fusion algorithm, enhancing positioning stability and accuracy. Furthermore, it can be equipped with dedicated motion sensor fusion chips (such as the VectorNav VN-100 series) for collaborative processing, effectively suppressing IMU drift errors and improving the stability of attitude estimation in complex terrain. Simultaneously, it ensures that the robotic arm's movements are based on real-time, high-precision spatial attitude, reducing picking errors. Based on the fused visual and inertial data, it predicts the robot's six-degree-of-freedom attitude (position, velocity, and attitude quaternions) in real time, providing accurate spatial positioning information for robotic arm control and path planning.
[0075] The defect detection module is electrically connected to the robotic arm control module. It is used to detect and perform preliminary classification of internal defects in pineapple fruits based on multispectral structured light data, and then send the classification results to the robotic arm control module.
[0076] The defect detection module can be implemented using the CPU / GPU resources in the attitude prediction module. It can use a pre-trained multimodal deep learning model to analyze the fusion features of multispectral and structured light, so as to achieve non-destructive detection and primary classification of internal defects in pineapple fruits, identify defects such as hollowness, rot, and insect infestation, thereby achieving high-accuracy internal quality detection of fruits, improving the quality control of harvested fruits, and reducing the cost of manual sorting in the later stage.
[0077] The robotic arm control module is used to dynamically adjust the picking strategy and the robot's robotic arm movements based on the posture prediction and classification results.
[0078] The robotic arm control module can be constructed using an ABB IRC5 series industrial robotic arm controller, a CX2040 series motion controller, and an NI myRIO series real-time control system. Based on posture prediction results and defect detection classification results, it dynamically plans the robotic arm's motion trajectory and picking sequence to perform efficient and precise fruit picking operations. This improves the accuracy and flexibility of the robotic arm's movements, reduces the damage rate of the picking machinery, and enhances overall picking efficiency and automation by dynamically adjusting the picking strategy.
[0079] The path planning module is electrically connected to the drive control module. It is used to generate the robot's motion path in complex terrain based on the robot's positioning data and preset 3D environmental data, and then send it to the drive control module.
[0080] The path planning module can also be implemented based on the CPU / GPU resources in the attitude prediction module. It can also be paired with a path planning software framework that supports GPU acceleration (such as ROS Navigation Stack and MoveIt!). Based on the robot's real-time localization and the pre-stored 3D terrain model, it can use a heuristic search algorithm to plan obstacle avoidance paths, ensuring that the robot can move safely and efficiently in complex terrain. This enables dynamic obstacle avoidance and adaptive terrain walking, improving operational safety, optimizing movement paths, shortening harvesting time, and reducing energy consumption.
[0081] The drive control module is used to control the robot to move according to the received motion path.
[0082] The drive control module can use Maxon EPOS4 series wheel / track drive controllers, combined with low-latency motor drivers and encoder feedback systems, to drive the robot chassis to perform precise movements according to the instructions of the path planning module, achieve path tracking and area coverage, and realize smooth and precise robot motion control to ensure that the robot can efficiently complete the coverage of the picking area and the continuity of the operation.
[0083] When fusing visual and inertial measurement data, the posture prediction module uses a combination of extended Kalman filtering and nonlinear optimization methods to estimate the robot's posture. The specific logic is as follows:
[0084] Time synchronization and spatial calibration are performed on the visual data and inertial measurement data collected by the robot;
[0085] Robot pose prediction is performed based on extended Kalman filtering, and pose update is performed by combining visual data.
[0086] The robot's historical poses are globally optimized using a nonlinear optimization method, thereby reducing accumulated errors.
[0087] The final output is a prediction of the robot's real-time position and attitude in complex terrain.
[0088] The robot's pose prediction is defined in the following vector form:
[0089]
[0090] In the formula Let p(t), v(t), and q(t) represent the robot's pose vector at time t, respectively, and let b be the robot's position, velocity, and pose quaternions at time t. a (t), b g(t) represents the biases of acceleration and angular velocity, respectively. These two biases can be obtained from the accelerometer and gyroscope in the inertial measurement unit. , represent the accelerometer and gyroscope, respectively. , represent the systematic error components present when measuring acceleration and angular velocity, respectively. These are manifested as constant or slowly changing offsets in the sensor output. These two biases are key parameters in the error model of the inertial measurement unit (i.e., IMU). Through real-time estimation and correction using state estimation algorithms (such as extended Kalman filtering), the impact of sensor errors on robot spatial positioning and attitude estimation can be reduced, ensuring positioning accuracy and system stability.
[0091] Attitude quaternions are a mathematical tool used to represent the rotational state of objects (such as robots or robotic arms) in three-dimensional space. They describe the rotational relationship from a reference coordinate system (usually the ground or world coordinate system) to the object's own coordinate system. Compared to traditional Euler angles or rotation matrices, quaternions are more computationally efficient and avoid gimbal lock.
[0092] Problems such as Gimbal Lock are pose representation methods widely used in modern robotics and aerospace.
[0093] Specifically, a pose quaternion can be represented as:
[0094]
[0095] The q here w q represents the scalar part of the real part. x q y q z The vector part representing the imaginary part, i, j, k are different from the subscripts in the following text, and here they represent the imaginary unit of the imaginary part.
[0096] Suppose that the rotation axis U = [u] around a certain unit is... x ,u y ,u z ] T When rotating by an angle ∈, the corresponding attitude quaternion can be represented as:
[0097]
[0098] Attitude quaternions can represent the robot's current spatial attitude (i.e., rotational state) to ensure that the robot can accurately perceive its own orientation in complex terrain, thereby providing reliable attitude information for robotic arm control and path planning.
[0099] When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing a cost function, the expression of which is shown below:
[0100]
[0101] In the formula, r(IMU,t) represents the IMU pre-integration residual at time t, r(vision,t) represents the visual measurement residual at time t, and ∑IMU,t and ∑vision,t represent the covariance matrices of the IMU pre-integration residual and the visual measurement residual at time t, respectively. This represents the weighted Euclidean norm.
[0102] As can be seen from the expression of the cost function, the cost function aims to estimate the robot's state vector so that the residual (error) of combining IMU pre-integrated data and visual measurement data is minimized, thereby achieving optimal spatial localization and attitude estimation.
[0103] Specifically, IMU data can provide continuous instantaneous acceleration and angular velocity information, which can be used to calculate motion trajectories through inertial navigation. However, its integration process can accumulate errors, especially when there is noise and bias. Visual data, on the other hand, provides relative pose measurement (such as visual odometry), which can achieve pose constraints through image feature matching and effectively suppress the accumulation of IMU errors. However, it is affected by occlusion and lighting. The cost function balances the residuals of these two types of data to make the positioning results take into account both the dynamic continuity of IMU and the geometric constraints of visual measurement, thereby improving robustness and accuracy.
[0104] The IMU (Inertial Measurement Unit) uses accelerometers and gyroscopes to calculate the predicted robot state, and the difference between this prediction and the current state of the optimized variables is the residual. The IMU residual reflects the degree to which the state variables interpret the inertial measurements; state estimates that do not conform to the IMU motion model will produce large residuals. The visual residual originates from the relative pose estimated by feature point matching between consecutive image frames, and the difference between this estimate and the pose at the corresponding time point in the optimized state. This residual constrains the robot's position and attitude, helping to correct the drift generated in the IMU integration. The covariance matrix of these two parameters is a statistical description of measurement noise; higher-precision sensors correspond to smaller covariances, and this residual is given greater weight during optimization to achieve a reasonable fusion of data reliability. Therefore, by minimizing the weighted sum of squared residuals, the algorithm, while considering the dynamic continuity of the IMU, fully utilizes the spatial constraints of visual measurements to obtain the robot state estimate that best matches multi-source observations. This ensures that the robot can obtain high-precision and robust pose and attitude information in complex terrain environments, thereby supporting subsequent precise picking and path planning by the robotic arm.
[0105] The defect detection module incorporates a multimodal deep learning model for detecting internal defects in pineapple fruits. The training logic for the multimodal deep learning model is as follows:
[0106] Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the types of internal defects were labeled.
[0107] Construct a fusion feature vector of multispectral-structured light and input it into a multimodal deep learning model for training;
[0108] Cross-entropy is used as the loss function, and the model parameters of the multimodal deep learning model are iteratively optimized by minimizing the loss function. During the iteration, forward propagation (calculating the output) and backpropagation (calculating the gradient and updating the parameters) are performed repeatedly until the loss function converges or the preset number of iterations is reached, at which point the model parameters are considered optimized.
[0109] This training step is to establish a labeled dataset that covers a variety of defect types (such as hollow, rotten, insect-eaten, etc.) to ensure the diversity and representativeness of the training samples.
[0110] in:
[0111] The expression for fusing feature vectors is:
[0112]
[0113] In the formula I(λ) represents the fused feature vector of the pixel at coordinates (x, y). i (x, y) represents wavelength λ i The time coordinate is the multispectral reflectance of the pixel at (x,y), D(x,y) represents the structured light depth value of the pixel at (x,y), k(x,y) represents the curvature feature of the pixel at (x,y), x and y represent the horizontal and vertical coordinates of the pixel in the image, respectively, the subscript i represents the index of the selected wavelength, and N is the total number of selected wavelengths. When selecting wavelengths, the bands that maximize the difference in reflectance between internal defects (such as decay and lesions) and healthy tissue should be selected to improve the sensitivity and accuracy of defect detection. For example, water (700nm~900nm), chlorophyll (670nm~680nm), and carotenoids (560nm~590nm) show significant differences in specific near-infrared and visible light bands, and these wavelengths should be given priority.
[0114] As can be seen from the expression of the fused feature vector, it is divided into three parts. The first part is multispectral reflectance, which is used to reflect the differences in the internal components (such as water, sugar, pigments, etc.) and surface state of the fruit. The second part is structured light depth feature, which is used to reflect the three-dimensional spatial position and morphological information of the pixel on the fruit surface, so as to reveal the spatial features such as depressions and expansions inside the fruit. The third part is curvature feature, which can be obtained through structured light depth map, and is used to describe the degree of curvature of the local surface of the fruit, that is, the rate of change of surface shape, so as to locate the defect boundary and identify the abnormality of the fruit surface.
[0115] When weighting the fused feature vector, a self-attention mechanism is introduced to extract the contextual latent state features of each pixel, and to calculate the dynamic weight of each vector element based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as follows:
[0116]
[0117] In the formula w ∈ (x,y) represents the dynamic weight of the ∈-th vector element in the fused feature vector, h(x,y) represents the contextual latent state feature of the pixel at coordinates (x,y), ∈ represents the index of the vector element, and f qz (·) represents the weight generation function. It can be understood that since the vector elements in the fused feature vector include N sets of multispectral reflectance, one set of structured light depth values, and one set of curvature features, the total number of its elements is N+2.
[0118] Contextual latent state features represent the contextual information state of the current pixel. They capture the surrounding environment and semantic information of the pixel, helping the weight generation function understand the importance of features at that location. Specifically, the semantic and spatial information of the pixel and its neighborhood can be extracted from the intermediate layer feature maps of a multimodal deep learning model, combined with global contextual information generated by self-attention mechanisms (Transformer, nonlocal networks, etc.), and finally obtained through feature fusion. This facilitates the identification of environmental factors such as illumination changes and occlusion, adjusts modal weights to reduce the contribution of affected modalities, and, combined with neighborhood defect distribution information, enhances the smoothness and consistency of modal weights within local continuous regions. This allows for dynamic adjustment of the sensitivity of different modalities to specific defect types, achieving more accurate defect localization.
[0119] The weight generation function can be designed as a parameterized neural network module; specifically, its expression can be set as:
[0120]
[0121] In the formula, A and B represent the weights and biases of the hidden layer, respectively, and a ∈ b ∈ Let represent the weights and biases of the ∈-th vector element in the output layer, respectively, and σ(·) represent the activation function, such as ReLU or LeakyReLU. This represents feature concatenation. This structure allows dynamic weights to capture the non-linear relationship between context and features, dynamically reflecting the environmental information at that location and the importance of the current feature.
[0122] The fused feature vector is then weighted using dynamic weights, calculated as follows:
[0123]
[0124] In the formula represents the weighted fused feature vector, and ⊙ represents the element-wise product;
[0125] The expression for the cross-entropy loss function is:
[0126]
[0127] In the formula, θ represents the model parameters to be trained, and y m , λ and λ represent the true label and model predicted label of the internal defect category of the sample, respectively. The subscript m represents the sample index, and λ represents the regularization hyperparameter. Representing the spatial gradient of the weights can smooth the weight distribution and improve the robustness of the model.
[0128] Cross-entropy measures the difference between the model's predicted probability distribution and the true label distribution. By continuously adjusting the parameter θ through optimization algorithms such as gradient descent, the predicted probability gradually approaches the true label, thereby improving the classification accuracy.
[0129] Specifically, the multimodal deep learning model here is a holistic machine learning framework or system designed to process and fuse multi-source heterogeneous data (such as multispectral image data, structured light depth information, etc.) to extract comprehensive features and complete the target task (such as internal defect detection and classification). The classifier can employ a neural network discriminant function, responsible for classifying or discriminating the fused features and outputting the defect category and the corresponding category probability.
[0130] The logic for using a trained multimodal deep learning model to perform preliminary classification of pineapple fruits is as follows:
[0131] Feature fusion was performed on the multispectral-structured light data of pineapple fruits collected in real time to obtain the fused feature vector of multispectral-structured light.
[0132] The fused feature vector is input into a multimodal deep learning model, which outputs the corresponding internal defect category and defect probability.
[0133] The defect score is calculated by multiplying the severity of each internal defect category by its probability. This score is then used to grade the quality of the pineapple fruit. The severity of each internal defect can be determined by expert experience, with higher scores indicating more severe defects. For example, hollowness has a severity score of 1, insect infestation a score of 2, and rot a score of 3. Higher defect scores indicate lower quality grades, while lower scores indicate higher quality grades.
[0134] In this step, by combining multispectral optical information and structured light geometric information, the shortcomings of single-modality methods are supplemented, significantly improving classification accuracy, detection precision and robustness. At the same time, the defect score is calculated by probability-weighted severity, avoiding the binary judgment of simple hard classification, reflecting the severity and confidence of the defect. Multiple categories of defects and their corresponding weights can be flexibly set to meet the different fruit quality management needs.
[0135] The logic of the robotic arm control module dynamically adjusting the picking strategy and the robot's robotic arm movements based on posture prediction and classification results is as follows:
[0136] The position and orientation of the robotic arm base are predicted based on the robot's real-time position and orientation.
[0137] The robotic arm's harvesting path is planned according to the quality grade of the pineapple fruit, with priority given to harvesting pineapple fruits of higher quality grade.
[0138] After each harvest, the harvesting order is dynamically adjusted and the robotic arm's harvesting path is updated until all pineapple fruits in the area have been harvested.
[0139] It is understandable that the robotic arm is mounted on a fixed rigid body structure of the robot (usually the robot chassis, body, or frame), therefore its pose has a fixed rigid transformation relationship with the robot's pose. Let the fixed rigid body transformation of the robotic arm base relative to the robot's body coordinate system be:
[0140]
[0141] In the formula T b / r Let R represent the transformation matrix. b / r p represents the rotation matrix of the robotic arm base relative to the robot coordinate system. b / r R represents the translation vector of the robotic arm base relative to the robot coordinate system. b / r p b / r All are fixed constants;
[0142] Then convert the robot's position and orientation quaternions into homogeneous transformation matrices:
[0143]
[0144] Where T r (t) denotes the homogeneous transformation matrix, R r (t) represents the rotation matrix obtained by transforming the attitude quaternion q(t);
[0145] According to the rigid body transformation chain, the attitude vector of the robotic arm base in the environment coordinate system can be expressed as:
[0146] T b (t)=T r(t)·T b / r
[0147] Right now:
[0148]
[0149] If we let:
[0150] R b (t)=R r (t)·R b / r
[0151] p b (t)=R r (t)·p b / r +p(t)
[0152] The attitude vector of the robotic arm base can then be expressed as:
[0153]
[0154] Here R b (t) is a rotation matrix representing the rotational attitude of the robotic arm base relative to the global environment coordinate system at time t, i.e., the orientation and attitude direction of the robotic arm base; p b (t) is a three-dimensional vector that represents the position coordinates of the robotic arm base in the global environment coordinate system at time t, that is, the position of the spatial point where the robotic arm base is located.
[0155] It can be seen that the attitude vector of the robotic arm base is the combination of the position and attitude in the robot's attitude vector and the fixed transformation of the base relative to the robot. In other words, the robot's global pose determines the spatial position of the robotic arm base, while the fixed geometric parameters of the robotic arm base relative to the robot body determine the rigid body transformation relationship between the two.
[0156] If we further divide the location information and quality grade information of the pineapple fruit into a set S, then this set can be represented as:
[0157] S={(Pg j ,L j )|j=1,2,…,M}
[0158] In the formula, Pg j L represents the position of the j-th pineapple fruit. j Let M represent the defect score of the j-th pineapple fruit, M represent the total number of pineapple fruits to be harvested in the area, and j represent the index of the pineapple fruit.
[0159] For set S, according to defect score L j The pineapple fruits are sorted in ascending order, and the sorted set is denoted as S′. The pineapple fruits are arranged from highest to lowest grade to facilitate the priority picking of higher quality fruits.
[0160] For a sorted set S, the picking path of the robotic arm can be represented as a sequence Q, where fruits are picked in order.
[0161] Q = [Pg] j1 ,Pg j2 ,…,Pg jq ,…,Pg jm ]
[0162] In the formula, Pg jq This represents the position of the q-th picking target, with the subscript jq indicating the index of the picking target, obtained from the sorted set S′.
[0163] When updating and optimizing the picking route, the goal is to minimize the total route length; therefore, the optimization function can be expressed as:
[0164]
[0165] This formula expresses the minimum sum of distances between two adjacent picking points along the path, aiming to reduce the total distance traveled by the robotic arm and improve efficiency.
[0166] In this step, by combining the robot's global pose vector with the rigid body transformation of the robotic arm base, the real-time pose of the robotic arm base in the environment is accurately obtained, ensuring that the robotic arm's movements are based on accurate spatial references, reducing motion errors and collision risks. Then, the fruit defect score is used to intelligently sort the picking targets. Compared with the traditional robotic arm positioning method for picking, this method of using fruit quality as a reference to correct the robotic arm's positioning and motion trajectory can achieve "prioritizing the picking of high-quality fruits", thereby improving the overall quality of the produced fruits, reducing the picking of low-quality fruits, and reducing resource waste.
[0167] The logic for planning the robot's motion path is as follows:
[0168] Identify the robot's traversable areas during its movement based on the robot's location data and a pre-set terrain model;
[0169] The algorithm uses a heuristic search algorithm to plan the shortest path from the current location to the next location, while avoiding obstacles and steep slopes.
[0170] Generate a continuous motion trajectory until all pineapple fruits in all areas have been harvested.
[0171] By combining the robot's real-time location data with a pre-built 3D terrain model, the robot can accurately identify navigable areas in complex terrain, avoiding unnavigable paths caused by blind planning and improving the safety and practicality of path planning.
[0172] In this embodiment, the multimodal scheme of the present invention is compared with two sets of traditional single-modal schemes (single IMU scheme and single vision scheme). Pineapple fruits are harvested from 20 regions respectively. After each region is harvested, parameters such as position error, posture error, defect accuracy, and harvesting path are counted. The total parameters in the harvesting process are as follows:
[0173] Table 1: Harvesting Parameters for Multimodal Schemes
[0174]
[0175]
[0176] Table 2: Parameters Acquired by Single-Mode (IMU) Scheme
[0177]
[0178]
[0179] Table 3: Acquisition Parameters for Single-Modal (Vision) Scheme
[0180]
[0181]
[0182] Understandably, the single IMU solution cannot detect defects in the fruit, so the defect detection accuracy is calculated manually. The single vision solution uses a traditional visual recognition algorithm (YoLoV5 model) to detect whether there are defects. Since the single multispectral structured light solution only involves fruit defect recognition and does not involve kinematic directions such as position estimation and attitude estimation, no control solution is set up.
[0183] Please refer to Figure 3 After merging all the parameters in the table above, the minimax normalization method is applied to eliminate dimensions. Specifically, for the parameters where smaller values are better (i.e., position estimation error, attitude estimation error, and total length of the picking path), the normalization formula is:
[0184]
[0185] For indicators where larger values are better, the normalization formula is:
[0186]
[0187] After normalization, the average of the data in the 20 regions can be plotted.
[0188] The advantage of doing this is that it not only eliminates the dimensions, but also eliminates the differences between the decision logic. That is, the closer the normalized value of all parameters is to 1, the better the performance, and the closer it is to 0, the worse the performance. In this way, a bar chart can be built to intuitively show the differences between the three different schemes.
[0189] As can be seen from the table data and bar charts, although the IMU (Inertial Measurement Unit) can continuously output acceleration and angular velocity information, the error accumulates rapidly over time (drift phenomenon) when its data is integrated for positioning. This is especially true in complex terrain and long-term operation scenarios, where positioning errors are often significant. Relying on IMU data and lacking auxiliary correction from external sensors such as vision and GPS, it cannot directly perceive its environment. The lack of visual information leads to weak path planning and obstacle avoidance capabilities, easily resulting in unreasonable path planning and obstacle avoidance failures. Furthermore, it cannot detect fruit defects. Therefore, the performance of its various parameters is very... The low performance of traditional vision-based solutions makes them unsuitable for practical applications. In contrast, single-vision solutions perform better because vision cameras can acquire rich environmental information, including color images and depth information (RGB-D), supporting more accurate feature extraction and environmental modeling. However, they lack stability and real-time performance in dynamic or complex terrains, affecting the continuity of robotic arm control and path planning. The multimodal solution of this invention has superior performance in all parameters. By fusing various heterogeneous data, it can dynamically adjust the robotic arm's trajectory and picking strategy, prioritizing the picking of high-quality fruits, reducing the risk of mechanical damage, and improving production efficiency.
[0190] A spatial heterogeneous visual localization method for a pineapple harvesting robot, applicable to the aforementioned localization system, includes the following steps:
[0191] S1: Collect multimodal visual data and inertial measurement data of the robot in the complex terrain environment, and at the same time collect multispectral and structured light data of the target fruit;
[0192] S2: The robot achieves spatial localization and attitude prediction through a vision-inertial fusion algorithm. At the same time, it uses multispectral and structured light data to detect internal defects in fruits in real time and perform primary classification based on a multimodal deep learning model.
[0193] S3: Control the picking path of the robotic arm based on the robot's position, posture and the quality grade of the pineapple fruit to complete the pineapple picking in the area where the robot is located;
[0194] S4: Based on the robot's location data and the preset terrain model, plan the robot's movement path and repeat steps S1 to S3 in each area until all pineapple fruits in all areas have been harvested.
[0195] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0196] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0197] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A spatial heterogeneous visual positioning system for a pineapple harvesting robot, characterized in that, Specifically, it includes: The data acquisition module is electrically connected to the attitude prediction module and the path planning module. It is used to collect the robot's visual data, inertial measurement data, position data and multispectral-structured light data of pineapple fruit in complex terrain, and send them to the attitude prediction module. The attitude prediction module is electrically connected to the robotic arm control module. It is used to predict the robot's attitude based on visual data and inertial measurement data, and send the attitude prediction results to the robotic arm control module. The defect detection module is electrically connected to the robotic arm control module. It is used to detect and perform preliminary classification of internal defects of pineapple fruits based on multispectral structured light data, and send the classification results to the robotic arm control module. The posture prediction module estimates the robot's posture by combining extended Kalman filtering and nonlinear optimization methods when fusing visual and inertial measurement data. The specific logic is as follows: Time synchronization and spatial calibration are performed on the visual data and inertial measurement data collected by the robot; Robot pose prediction is performed based on extended Kalman filtering, and pose update is performed by combining visual data. The robot's historical poses are globally optimized using a nonlinear optimization method, thereby reducing accumulated errors. The final output is a real-time prediction of the robot's position and attitude in complex terrain. A robotic arm control module is used to dynamically adjust the picking strategy and the robotic arm movements of the robot based on the posture prediction results and classification results. The logic of the robotic arm control module dynamically adjusting the picking strategy and the robot's robotic arm movements based on the posture prediction and classification results is as follows: The position and orientation of the robotic arm base are predicted based on the robot's real-time position and orientation. The robotic arm's harvesting path is planned according to the quality grade of the pineapple fruit, with priority given to harvesting pineapple fruits of higher quality grade. After each harvest, the harvesting order is dynamically adjusted and the robotic arm harvesting path is updated until all pineapple fruits in the area have been harvested. The path planning module is electrically connected to the drive control module and is used to generate the robot's motion path in complex terrain based on the robot's positioning data and preset environmental 3D data, and send it to the drive control module. The planning logic for the robot's motion path is as follows: Identify the robot's traversable areas during its movement based on the robot's location data and a pre-set terrain model; The algorithm uses a heuristic search algorithm to plan the shortest path from the current location to the next location, while avoiding obstacles and steep slopes. Generate a continuous motion trajectory until all pineapple fruits in all areas have been harvested; A drive control module is used to control the robot to move according to the received motion path.
2. The spatial heterogeneous visual positioning system for a pineapple harvesting robot according to claim 1, characterized in that: The visual data is an RGB-D image of the robot's location; the inertial measurement data includes the robot's velocity, acceleration, and angular velocity during movement; the multispectral-structured light data includes multispectral images and structured light depth maps, and covers the near-infrared and short-wave infrared bands.
3. The spatial heterogeneous visual positioning system for a pineapple harvesting robot according to claim 2, characterized in that: The robot's pose prediction is defined in the following vector form: In the formula Indicates that the robot is The pose vector at time step, , , These represent the robot in The position, velocity, and attitude quaternions at each moment. , These represent the biases for acceleration and angular velocity, respectively. When fusing the robot's visual data and inertial measurement data, this is achieved by minimizing a cost function, the expression of which is shown below: In the formula express The IMU pre-integral residual at time step [time]. express Visual measurement residuals at time points , They represent The covariance matrix of the IMU pre-integration residual and the visual measurement residual at time step [time]. This represents the weighted Euclidean norm.
4. The spatial heterogeneous visual positioning system for a pineapple harvesting robot according to claim 3, characterized in that: The defect detection module incorporates a multimodal deep learning model for detecting internal defects in pineapple fruits. The training logic of the multimodal deep learning model is as follows: Several groups of pineapple fruits with internal defects were used as samples. Multispectral images and structured light depth maps of the samples were collected, and the types of internal defects were labeled. Construct a fusion feature vector of multispectral-structured light, perform weighted processing on it, and then input it into a multimodal deep learning model for training; Cross-entropy is used as the loss function, and the model parameters of the multimodal deep learning model are iterated repeatedly until the optimization is completed, with the goal of minimizing the loss function. in: The expression for the fused feature vectors is: In the formula Indicates coordinates as The fused feature vector of the pixel. Indicates wavelength as Time coordinates are Multispectral reflectance of the pixel Indicates coordinates as The structured light depth value of the pixel. Indicates coordinates as The curvature features of the pixels. , These represent the x and y coordinates of pixels in the image, respectively, and the subscripts are... Indicates the index of the selected wavelength. This represents the total number of wavelengths selected. When weighting the fused feature vector, a self-attention mechanism is introduced to extract the contextual latent state features of each pixel, and to calculate the dynamic weight of each vector element based on the contextual latent state features and the fused feature vector. The calculation formula is expressed as follows: In the formula Represents the th element in the fused feature vector. The dynamic weights of each vector element Indicates coordinates as The contextual hidden state features of a pixel. Indicates the index of a vector element. This represents the weight generation function; The fused feature vector is then weighted using dynamic weights, calculated as follows: In the formula This represents the fused feature vector after weighted processing. Represents element-wise product; The expression for the cross-entropy loss function is: In the formula This represents the parameters of the model to be trained. , These represent the true labels and model-predicted labels for the internal defect categories of the sample, respectively, with subscripts... Indicates the index of the sample. This represents the regularization hyperparameter. Represents the spatial gradient of the weights.
5. The spatial heterogeneous visual positioning system for a pineapple harvesting robot according to claim 4, characterized in that: The logic for using a trained multimodal deep learning model to perform preliminary classification of pineapple fruits is as follows: Feature fusion was performed on the multispectral-structured light data of pineapple fruits collected in real time to obtain the fused feature vector of multispectral-structured light. The fused feature vectors are weighted and then input into a multimodal deep learning model, which outputs the corresponding internal defect category and defect probability. The product of the severity of the internal defect category and the defect probability is used as the defect score, and the quality of the pineapple fruit is graded based on the defect score.
6. A spatial heterogeneous visual positioning method for a pineapple harvesting robot, characterized in that: The positioning method is applicable to the positioning system according to any one of claims 1-5, and the specific steps include: S1: Collect multimodal visual data and inertial measurement data of the robot in the complex terrain environment, and at the same time collect multispectral and structured light data of the target fruit; S2: The robot achieves spatial localization and attitude prediction through a vision-inertial fusion algorithm. At the same time, it uses multispectral and structured light data to detect internal defects in fruits in real time and perform primary classification based on a multimodal deep learning model. S3: Control the picking path of the robotic arm based on the robot's position, posture and the quality grade of the pineapple fruit to complete the pineapple picking in the area where the robot is located; S4: Based on the robot's location data and the preset terrain model, plan the robot's movement path and repeat steps S1 to S3 in each area until all pineapple fruits in all areas have been harvested.
Citation Information
Patent Citations
Spatial heterogeneous visual positioning system of tomato picking robot
CN119772891A
Cooperative operation method, system and platform for guiding robot to pick and convey fruits
CN117546681A
Cotton picking robot and picking navigation method thereof
CN120167228A