Low damage picking method, robot with multi-modal perception and cross-view imitation

CN122807826APending Publication Date: 2026-09-25SHENZHEN WEIXIA ROBOT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611243417.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-17
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明所要解决的问题是强光/背光与枝叶遮挡导致检测不稳定、遮挡场景下目标定位精度不足容易引发抓取失败、规划层对柔性障碍建模不足、动作策略依赖人工规则缺乏学习与泛化能力等至少一个技术问题

Benefits of technology

所述处理器,用于当执行所述计算机程序时,实现如第一方面所述的多模态感知以及跨视角模仿的低损伤采摘方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807826A_ABST
    Figure CN122807826A_ABST
Patent Text Reader

Abstract

The application provides a low-damage picking method and robot for multi-modal perception and cross-view imitation, and relates to the technical field of intelligent robots. The method comprises the following steps: after time synchronization calibration is performed on multi-source picking observation data, detection segmentation and target enhancement are executed, maturity evaluation, three-dimensional positioning and uncertainty indicators are combined, and target list data for picking are output; based on the target list data and reachable parking area data and reachable posture data of the robot, comprehensive cost evaluation is performed, and the comprehensive cost of each target is obtained; an optimal parking point is determined and an approaching trajectory is generated; after cooperative planning is performed to reach the optimal parking point, the robot is driven to move to the target position according to the approaching trajectory, a strategy model obtained through cross-view imitation learning training is used to obtain a motion block sequence, the motion block sequence is filtered through safety constraints, and the robot is driven to complete picking actions in a force control mode. The application can reduce the damage rate and improve the picking success rate and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent robot technology, and more specifically, to a low-damage harvesting method and robot using multimodal perception and cross-viewpoint imitation. Background Technology

[0002] With the advancement of smart agriculture and large-scale planting, the labor shortage and rising costs in the fruit and vegetable harvesting process have become increasingly prominent. Compared with large fruits such as apples and citrus, small fruits such as blueberries are characterized by their small diameter, dense distribution, and color that is easily confused with the leaf background. In addition, multiple ripening stages coexist on the same branch. At the same time, blueberry stems are fragile, and excessive clamping force or inappropriate detachment methods can cause damage to the fruit skin, juice leakage, and a decline in quality.

[0003] In related technologies, harvesting systems often employ a modular "perception-planning-execution" pipeline: first, the fruit is located based on image recognition; then, the trajectory of the robotic arm is generated; and finally, the control module executes the process. This method can achieve certain results in experimental greenhouses or simplified environments; however, in real orchard environments, it faces four main challenges when harvesting small fruits such as blueberries: unstable detection due to strong light / backlighting and foliage obstruction; insufficient target localization accuracy in obstructed scenarios, easily leading to grasping failures; insufficient modeling of flexible obstacles in the planning layer; and a lack of learning and generalization capabilities due to reliance on manually defined rules for action strategies. Summary of the Invention

[0004] The problem to be solved by this invention is at least one of the following technical issues: unstable detection caused by strong light / backlight and foliage occlusion; insufficient target positioning accuracy in occluded scenes that easily leads to grasping failure; insufficient modeling of flexible obstacles in the planning layer; and lack of learning and generalization ability for action strategies that rely on manual rules.

[0005] To address the aforementioned issues, in a first aspect, the present invention provides a low-damage harvesting method based on multimodal perception and cross-viewpoint imitation, comprising: detecting and segmenting multi-source harvesting observation data that has been time-synchronized and calibrated; outputting individual perception data that has undergone target enhancement and occlusion robustness optimization processing; assessing the maturity and three-dimensional localization of targets based on the individual perception data; and determining uncertainty indicators to comprehensively output a target list data for harvesting. Based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, a comprehensive cost assessment is conducted on the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost to obtain the comprehensive cost of each target. The current objective is determined based on the comprehensive cost, and the operation of the robotic arm and the movement of the chassis are coordinated based on the map accessibility data under the current objective to determine the optimal stopping point and generate an approach trajectory. When the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to the target location according to the approach trajectory, the strategy model trained through cross-view imitation learning is input with the target list data and the picking task to obtain the action block sequence and perform safety constraint filtering, so as to drive the robotic arm to complete the picking action through damage prevention control. The low-damage picking method based on multimodal perception and cross-viewpoint imitation provided by this invention performs time-synchronized calibration, target enhancement, and occlusion robustness optimization on multi-source picking observation data. Combined with maturity assessment, 3D localization, and the joint output of uncertainty indicators, it can stably output high-quality target list data even under conditions of varying lighting and dense occlusion, fundamentally solving the problem of grasping failure caused by detection instability and single-viewpoint depth estimation errors. Based on obtaining a reliable target list, the system comprehensively evaluates picking benefits, collision risks, accessibility, detection uncertainty, and action switching costs by incorporating them into a weighted comprehensive cost function. This allows the system to fully weigh "whether to pick" and "whether it can pick" during the sorting stage. Furthermore, after determining the current target based on the comprehensive cost, the system uses map accessibility data to guide the robotic arm operation and chassis movement. The system actively performs collaborative planning, proactively searching for and relocating the optimal stopping point when the current stopping point is unreachable or too risky. This effectively reduces frequent chassis movement and large swings of the robotic arm during the planning phase, jointly addressing issues such as insufficient modeling of flexible obstacles, unreachable paths, or frequent collisions in the planning layer. Based on the completed planning, the strategy model trained through cross-perspective imitation learning maps the target list data and picking tasks into a sequence of action blocks. After being filtered by safety constraints, the picking is executed in a force-controlled manner. This allows the system to learn refined picking skills from human demonstrations and ensures the physical feasibility of the output actions, solving the problem of action strategies relying on manual rules and lacking generalization ability. At the same time, combined with the damage-prevention force-controlled execution method, it achieves a low-damage picking action sequence of "stabilize first, then detach, and then release slowly," significantly reducing fruit crushing and drop damage.

[0006] Secondly, the present invention also provides a robot that employs a low-damage picking method using multimodal perception and cross-viewpoint imitation as described in any of the preceding claims.

[0007] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the low-damage picking method for multimodal sensing and cross-viewpoint imitation as described in the first aspect.

[0008] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the low-damage picking method for multimodal sensing and cross-view imitation as described in the first aspect.

[0009] The robot, electronic device, and computer-readable storage medium provided by this invention have the same beneficial effects as the low-damage harvesting method of multimodal perception and cross-view imitation compared to the prior art, and will not be repeated here. Attached Figure Description

[0010] Figure 1 A flowchart illustrating a low-damage harvesting method using multimodal sensing and cross-viewpoint mimicry according to an embodiment of the present invention is shown. Figure 2 A schematic diagram of the robot's architecture in an embodiment of the present invention is shown; Figure 3 A schematic diagram of the robot's autonomous harvesting operation process is shown in an embodiment of the present invention; Figure 4 This diagram illustrates the multimodal fusion maturity assessment and 3D localization process in an embodiment of the present invention. Figure 5 A schematic diagram of the operational risk control and online quality assessment framework in an embodiment of the present invention is shown; Figure 6 This diagram illustrates a closed-loop system for cross-perspective imitation learning training, security, and deployment in an embodiment of the present invention. Figure 7 This illustration shows a data-model-deployment closed loop driven by digital twins in an embodiment of the present invention. Figure 8 This diagram illustrates the task scheduling and data aggregation for multi-robot collaborative harvesting in an embodiment of the present invention. Figure 9 A schematic diagram of the movable body of the robot in an embodiment of the present invention is shown; Figure 10 A side view of the robot's structure in an embodiment of the present invention is shown; Figure 11 A schematic diagram of the robot's front view structure in an embodiment of the present invention is shown; Figure 12 A schematic diagram of the end effector of a robot in an embodiment of the present invention is shown; Figure 13 A three-dimensional structural schematic diagram of the robot in an embodiment of the present invention is shown; Figure 14 A schematic diagram of the module composition of the robot in an embodiment of the present invention is shown.

[0011] Explanation of reference numerals in the attached figures: 1. Lifting platform; 2. Emergency stop switch; 3. Movable body; 4. Four-wheel omnidirectional independent chassis; 5. LiDAR; 6. Touch screen; 7. Robotic arm body; 8. Torque sensor; 9. End effector; 10. Robotic arm flange; 11. Vision camera; 12. Flexible adsorption mechanism one; 13. Flexible adsorption mechanism two; 14. Flexible gripper; 15. Fruit guide groove; 16. Storage basket; 17. Storage frame; 18. Storage basket guide. Detailed Implementation

[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0013] It should be noted that relational terms such as "first" and "second" in this invention are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0014] In the description of this specification, references to terms such as "embodiment," "one embodiment," and "one implementation" indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or implementation is included in at least one embodiment or illustrative implementation of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or implementation. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or implementations.

[0015] Reference Figure 1 As shown, this embodiment of the invention proposes a low-damage harvesting method with multimodal sensing and cross-viewpoint mimicry, comprising: The multi-source harvesting observation data that has been time-synchronized and calibrated is detected and segmented, and individual perception data that has been optimized by target enhancement and occlusion robustness is output. Based on the individual perception data, the maturity of the target is assessed and three-dimensionally located, and uncertainty indicators are determined to comprehensively output a target list data for harvesting.

[0016] Specifically, through collaborative processing of multi-source harvesting and observation data, the system can extract the spatial location, maturity status, and reliability of the assessment of each fruit from the raw sensor data, forming a structured target list. For example, during orchard operations, the data collected by the sensing unit, after time synchronization and fusion, can output an information set containing multiple fruit location coordinates, maturity levels, and uncertainty scores for subsequent decision-making processes. Detection and segmentation are used to distinguish fruits from the background and determine their outlines in the image; maturity assessment is used to determine whether the fruit has reached the harvestable standard; 3D localization is used to convert the fruit positions in the image into spatial coordinates that the robot can execute; and the uncertainty index is used to quantify the reliability of the current sensing results, preventing the system from making incorrect decisions based on unreliable information.

[0017] Based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, a comprehensive cost assessment is conducted on the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost to obtain the comprehensive cost of each target.

[0018] Specifically, comprehensive cost assessment involves a unified quantitative comparison of multiple decision factors for each fruit in the target list to determine the optimal harvesting order. For example, it can comprehensively consider the fruit's maturity benefit value (fruits with higher maturity are more worthy of priority harvesting), the safety risks of harvesting from the current location (the possibility of collisions with surrounding branches or obstacles), the ease with which the robotic arm can reach the fruit (whether it can reach it without collision from the current docking position), the reliability of the current perception results (the lower the perception uncertainty, the more reliable it is), and the time and energy required to switch from the current posture to the fruit's posture (the lower the switching cost, the better for improving overall efficiency). These factors are weighted and summed to obtain a comprehensive score. The fruit with the lowest score is the most suitable target for priority harvesting under the current comprehensive conditions, thus achieving multi-objective optimization decision-making.

[0019] Based on the comprehensive cost, the current objective is determined, and based on the map accessibility data under the current objective, the operation of the robotic arm and the movement of the chassis are coordinated and planned to determine the optimal stopping point and generate an approach trajectory.

[0020] Specifically, collaborative planning addresses the question of "how to get the robot to the harvestable state in the best way" after the target is determined, seeking the optimal balance between chassis movement and robotic arm operation. For example, the system first queries the robotic arm's coverage capability at the current chassis position. If the current docking point allows direct harvesting without safety hazards, the robotic arm's movement path is directly generated. If the current position is unsuitable (the robotic arm cannot reach the target) or poses a high harvesting risk, multiple candidate positions are searched through a pre-built accessibility map. The system comprehensively evaluates the number of targets each candidate position can cover, the distance and time cost of moving from the current position, and the difficulty of turning at that position. The optimal position is selected, and the chassis is driven to the position. The robotic arm path is then planned, thereby achieving joint optimization of chassis movement and robotic arm operation, avoiding frequent movements or large arm swings.

[0021] When the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, the strategy model trained by cross-view imitation learning is input with the target list data and the picking task to obtain the action block sequence and perform safety constraint filtering, so as to drive the robotic arm to complete the picking action through the damage prevention control method.

[0022] Specifically, the cross-perspective imitation learning strategy model is responsible for translating high-level task instructions into specific low-level action sequences. For example, the model pre-learns human picking demonstrations from different perspectives (such as the operator's head-mounted view, the observer's handheld view, and the fixed view of the scene), combined with language annotations during the demonstrations (such as "prioritize picking ripe fruit" and "avoid branches"), to establish a mapping relationship from scene perception to action output. During deployment and runtime, the model outputs a sequence of action blocks containing end-effector pose increments and control instructions such as gripper opening and closing and suction switches, based on the currently input target list data and the language instructions for the picking task. This output must be verified by a safety constraint layer (checking whether it exceeds joint limits, whether the speed and acceleration exceed safety limits, whether the end-effector force is within a safe range, and whether there is a risk of collision with surrounding obstacles). Only after the verification is passed can the robot be driven to perform the complete picking action from approaching, grasping, detaching, to placing, avoiding the model outputting infeasible and dangerous actions.

[0023] In practical application, this embodiment performs time synchronization calibration, target enhancement, and occlusion robustness optimization on multi-source harvesting observation data. Combined with maturity assessment, 3D positioning, and the joint output of uncertainty indicators, it ensures stable output of high-quality target list data even under varying lighting conditions and dense occlusion. This fundamentally solves the problem of grasping failure caused by detection instability and single-view depth estimation errors. Based on a reliable target list, the system incorporates harvesting benefits, collision risk, accessibility, detection uncertainty, and action switching costs into a comprehensive cost function for weighted evaluation. This allows the system to fully weigh "whether to harvest" and "whether it can harvest" during the sorting stage. Furthermore, after determining the current target based on the comprehensive cost, the system collaboratively plans the robotic arm operation and chassis movement based on map accessibility data. When the current stopping point is unreachable or the risk is too high, the system actively searches for the optimal stopping point and relocates it. This allows the system to effectively reduce frequent chassis movement and large swings of the robotic arm during the planning phase, jointly solving the problems of insufficient modeling of flexible obstacles, unreachable paths, or frequent collisions in the planning layer. Based on the completion of planning, the strategy model obtained through cross-view imitation learning maps the target list data and picking tasks into a sequence of action blocks. After being filtered by safety constraints, the picking is executed in a force-controlled manner. This allows the system to learn fine picking skills from human demonstrations and ensure the physical feasibility of the output actions, solving the problem of action strategies relying on manual rules and lacking generalization ability. At the same time, combined with the damage prevention force control execution method, a low-damage picking action sequence of "stabilize first, then detach, and then release slowly" is achieved, significantly reducing fruit crushing and drop damage.

[0024] Compared with related technologies, this invention, through the coordinated operation of the above four steps, addresses four core pain points in real orchard environments: unstable detection, inaccurate positioning, insufficient planning, and stiff movements. It achieves corresponding effects of "more stable detection, more accurate positioning, better path, smoother movements, and learnable experience," forming a complete closed loop from perception, decision-making, planning to execution, and jointly realizing the core goal of "reducing damage rate and improving harvesting success rate and efficiency."

[0025] It should be noted that, as Figure 2As shown, the harvesting method proposed in this invention can be deployed on a robot with basic autonomous operation capabilities. In one implementation architecture, the robot includes at least: a mobile carrier platform for moving between rows in the orchard and carrying operating components; a multi-joint robotic arm mounted on the mobile carrier platform for driving an end effector to move in space; an end effector located at the end of the robotic arm, having at least gripping and releasing functions; a sensing unit, including at least a vision sensor capable of acquiring color and depth images; and a computing and control unit, communicatively connected to the above components, for executing the data processing, decision-making, and motion control steps in the harvesting method described in various embodiments of this invention, and equipped with a storage medium to save pre-constructed map data and pre-trained model parameters. Based on this, the robot can achieve fully autonomous harvesting operations from target perception and decision-making planning to motion execution.

[0026] like Figure 3 and Figure 4 As shown, in an optional embodiment of the present invention, the multi-source harvesting observation data includes at least RGB images, near-infrared images, and depth data; The process involves detecting and segmenting multi-source harvesting observation data that has undergone time synchronization and calibration, outputting individual perception data that has been optimized for target enhancement and occlusion robustness, assessing target maturity and performing 3D localization based on the individual perception data, and determining uncertainty indicators to comprehensively output a target list for harvesting, including: A multi-scale feature fusion network for densely occluded target scenes is used to detect and segment the RGB image, and output the target bounding box and instance segmentation mask; Specifically, the time-synchronized RGB image is input into a lightweight detection network (such as an improved version of the YOLO series), and multi-scale feature fusion (FPN / PAN structure) and attention module (such as CBAM) are introduced to enhance the feature representation of small targets, and the target box and instance segmentation mask are output. During the training phase, data augmentation methods such as occlusion enhancement (random occlusion block), illumination perturbation (HSV random perturbation), motion blur (Gaussian kernel convolution) and background replacement are used.

[0027] The near-infrared image of the detected fruit ROI region is fused with the RGB image, and a maturity index is output through a classifier or regressor. Specifically, for each detected fruit ROI, the near-infrared channel image and the RGB channel image are stitched together along the channel dimension, or a two-branch feature fusion network is used. The result is then input into a lightweight classifier (such as MobileViT) or VisionTransformer, and the output is a maturity level (5-level discrete classification: green—color change—semi-ripe—ripe—overripe) or a continuous maturity index. ∈[0,1].

[0028] Based on the depth data, the pixel coordinates within the instance segmentation mask are transformed to three-dimensional coordinates in the robot base coordinate system to obtain the position of the target; Specifically, using depth maps or point cloud data, combined with the calibration parameters obtained in step two, the pixels within the segmentation mask are... Converted to 3D coordinates in the camera coordinate system, then... Transform to the robot base coordinate system, and take the centroid of the three-dimensional points in the masked area as the center position of the fruit. ;in For the camera intrinsic parameter matrix, For depth value, ( () represents pixel coordinates; then... Transform to the robot base coordinate system, where This is the transformation matrix from the camera to the end effector of the robotic arm. The transformation matrix from the robotic arm's end effector to the base; the centroid of the three-dimensional points in the masked region is taken as the target center position. , Indicates the first The three-dimensional position vector of the target in the robot's base coordinate system.

[0029] The uncertainty index is determined based on at least one of detection confidence, depth noise, occlusion ratio, and reprojection error; Specifically, uncertainty indicators The following formula is used for calculation: calculate; Overall detection confidence level (No. (Detection confidence score of each target), depth sensor noise) (Estimated by depth sensor noise model or local point cloud variance), occlusion ratio (The complement of the ratio of mask area to detection box area, representing the degree to which the target is occluded by branches and leaves) and reprojection error (The difference between the coordinates of the 3D point reprojected back to the image plane and the original pixel coordinates); Detection confidence complement ( Weights Based on the confidence calibration results of the detection network on the validation set, the depth noise was determined. weight The occlusion ratio is determined based on the calibration noise model of the depth sensor within the target working distance range. weight Based on experimental data on the impact of occlusion on 3D positioning accuracy, the reprojection error was determined. weight The weighting coefficients are determined based on the statistical distribution of hand-eye calibration accuracy and reprojection error. Each weighting coefficient satisfies... The specific values ​​are determined through calibration experiments under different lighting conditions, different degrees of shading, and different target distances, and can be adaptively adjusted according to the actual working environment of the orchard.

[0030] The target list data is determined based on the location of each target, the maturity index, and the uncertainty index.

[0031] Specifically, the location of each detected fruit target Maturity Index With uncertainty indicators The data is stored in association to form a target list for harvesting.

[0032] In practical applications, this embodiment, through multimodal fusion perception and uncertainty quantification, enables the system to stably detect and locate fruits in complex orchard environments with varying light levels and foliage obstruction; maturity assessment, combined with near-infrared channel information, can effectively distinguish between leaves of similar color and unripe fruits; the introduction of uncertainty indicators provides a quantitative basis for perceptual confidence in subsequent decision-making.

[0033] like Figure 4 As shown, as an optional embodiment of the present invention, the low-damage harvesting method of multimodal sensing and cross-viewpoint mimicry further includes: When the uncertainty index exceeds the first preset threshold, the chassis is moved and / or the robotic arm is swung and scanned to update the position of the target and the maturity index; if the uncertainty index still exceeds the first preset threshold after the update, the target is downweighted during the sorting stage.

[0034] Specifically, when considering comprehensive uncertainty indicators Exceeding the preset threshold When the system triggers a supplementary viewpoint action, including chassis micro-movement (preferably translation not exceeding 0.2m) and / or robotic arm swing scanning (preferably swing not exceeding 10°), it acquires new multi-view observation data and re-performs 3D positioning and maturity assessment; if the supplementary viewpoint is updated Still more If the overall cost of the target is increased during the overall cost ranking stage, its harvesting priority will be reduced.

[0035] In practical applications, the supplementary viewpoint mechanism reduces perception uncertainty by actively adjusting the observation pose, effectively solving the problems of depth estimation error and occlusion ambiguity under single viewpoint. The weighting process ensures that high-uncertainty targets will not fail to be harvested or damaged due to positioning errors. The system prioritizes targets with high perception credibility, improving the overall harvesting success rate and safety.

[0036] like Figure 3and Figure 5 As shown, in an optional embodiment of the present invention, the comprehensive cost assessment of harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost is performed based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data. The comprehensive cost of each target includes: The harvesting revenue is determined based on the maturity assessment results of each target; the collision risk is determined based on the relative positional relationship between each target and surrounding obstacles; the accessibility of each target is queried based on the accessibility map and the inverse accessibility map; and the action switching cost is determined based on the pose deviation between the current end pose and the poses of each target. Specifically, for candidate fruits Harvesting income According to the maturity index It is certain that the higher the maturity, the higher the yield, which can be further combined with the expected weight (based on fruit diameter). (Reverse estimation) and order strategy are comprehensively evaluated; collision risk By constructing a local obstacle map through voxelization of point cloud data, the distance between the target fruit and the nearest obstacle is calculated; the closer the distance, the higher the collision risk; accessibility score. Based on the current chassis parking position In reachability map If the current docking location is unreachable, a further query of the inverse reachability graph will be performed. Obtain candidate docking location information; cost of action switching The position is determined based on the positional deviation between the current end pose and the target picking pose. Specifically, it is a weighted distance estimate of the poses of both ends in joint space or Cartesian space, reflecting the time and energy cost required to switch from the current pose to the target pose.

[0037] The comprehensive cost of each objective is obtained by weighted summing of the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost.

[0038] Calculate the overall cost using the following formula: ; in, For the first The overall cost of achieving each goal; For harvesting revenue, at least based on the maturity index. Sure; This is a collision risk item; For reachability scores, based on the reachability graph With inverse reachability graph The query returned the result. As an uncertain indicator, based on the detection confidence level Depth noise Occlusion ratio With reprojection error ; Uncertainty Indicators The calculation has been described above; The cost of action switching is determined based on the pose deviation between the current end effector pose and the target pose. - These are the corresponding weighting coefficients used to balance the relative importance of each item in the overall cost. They can be adaptively adjusted according to the actual operating scenario (such as orchard varieties, lighting conditions, target density, etc.); the system selects the overall cost. The smallest target is used as the next harvesting target, and the parameters and weights are updated continuously during the execution process.

[0039] In practical applications, this embodiment uses a comprehensive cost function to quantify harvesting benefits, collision risks, accessibility, perception uncertainty, and action switching costs. This allows the system to simultaneously consider five dimensions at the decision-making level: "should we harvest?" (benefits), "can we harvest?" (accessibility and collision risks), "are we allowed to harvest?" (uncertainty), and "is it worthwhile to prioritize harvesting?" (switching costs). Through weighted summation, multi-objective comprehensive optimization is achieved, effectively improving the intelligence level and operational efficiency of harvesting decisions.

[0040] like Figure 5 As shown, in an optional embodiment of the present invention, the step of determining the current target based on the comprehensive cost, and coordinating the operation of the robotic arm and the movement of the chassis based on the map reachability data under the current target to determine the optimal docking point and generate an approach trajectory includes: If the current docking point is reachable from the target and the overall risk is not higher than the second preset threshold, the robotic arm trajectory planning is directly triggered to generate the approach trajectory from the current end pose to the target picking pose; Specifically, after calculating the comprehensive cost of each objective, the system sorts the objectives in ascending order according to the comprehensive cost. The smaller the comprehensive cost, the lower the comprehensive cost of picking the objective in the current state and the higher its priority. The objective with the smallest comprehensive cost is selected as the current objective to be picked and enters the collaborative planning stage: that is, for the selected objective, it is determined whether it is feasible and safe to pick the objective at the current docking point. After determining the target to be harvested, first query the accessibility information corresponding to the current chassis parking location: based on the current parking location In reachability map Query the 3D location of the target fruit Corresponding reachability score , To normalize to The continuous fraction of an interval is measured by the operability of the joint solution or the margin from the workspace boundary; if The collision risk is determined based on the relative positions of each target and obstacle at the current docking position, upon reaching the corresponding preset threshold. and uncertainty indicators If all are within the safe range, the robotic arm trajectory planning is directly triggered to generate an approach trajectory from the current end pose to the target picking pose. The trajectory planning uses sampling planning (such as RRT*) to sample and generate a collision-free path in joint space or Cartesian space, and then uses optimization planning (such as CHOMP or TrajOpt) to smooth and optimize the trajectory. The trajectory consists of five key path points: starting point, approach point, contact point, disengagement point and retreat point. The approach point is located 10-30mm in front of the target.

[0041] If the current docking point is unreachable from the target or the overall risk is higher than the second preset threshold, a set of candidate docking points is searched on the inverse reachability map. Each docking point in the candidate docking point set is comprehensively evaluated based on the number of reachable targets, average movement cost, driving distance, and turning difficulty indicators, and the optimal docking point is selected by weighting. The chassis is driven to move to the optimal docking point, and local fine positioning is performed by visual markers or ground structures after the chassis is in place. After positioning is complete, the robotic arm trajectory planning is triggered to generate the approach trajectory from the current end pose to the target picking pose.

[0042] Specifically, if at the current docking location Search get Below the corresponding preset threshold (i.e., the target is unreachable), or collision risk. Higher than the safety threshold, uncertainty index If any of the conditions exceeds the corresponding safety threshold, the system determines that the current docking point is unreachable or the overall risk is too high; at this time, the system will determine the target's three-dimensional position. In reverse reachability graph Query the set of candidate stops , It records the set of candidate chassis docking positions corresponding to any target pose in the workspace; for Each candidate stop The system calculates four indicators: the number of reachable targets (the number of fruits that can be covered at the docking point), the average movement cost (the estimated distance or time to move from the current docking point to the candidate point), the travel distance (the distance the chassis moves), and the steering difficulty (the attitude adjustment range the chassis needs to make at the docking point). By setting weight coefficients for each indicator, a weighted sum is calculated to obtain a comprehensive score. The docking point with the best comprehensive score is selected as the target docking point. Then, the chassis is driven to move to the optimal docking point. After the chassis is in place, local fine positioning is performed using visual markers (such as AprilTag) or ground structures to eliminate navigation accumulation errors. Once positioning is complete, the robotic arm trajectory planning is triggered. Using the same planning method as for the reachable case, an approach trajectory from the current end pose to the target picking pose is generated.

[0043] In practical applications, this embodiment utilizes reachability graphs. With inverse reachability graph Through offline pre-calculation and online querying, the system can quickly determine whether the current docking point is suitable for picking the current target, and efficiently search for the optimal docking point when needed. The collaborative planning strategy comprehensively considers multiple indicators such as the number of reachable targets, movement cost, driving distance and turning difficulty, avoiding frequent chassis movement and large swings of the robotic arm, reducing invalid actions and idle time, and improving the overall picking throughput per unit time while ensuring picking safety. For targets with good accessibility and controllable risks, the robotic arm planning is directly executed to shorten the single fruit picking cycle.

[0044] like Figure 5 , Figure 6 and Figure 7 As shown, in an optional embodiment of the present invention, when the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, the strategy model trained through cross-view imitation learning is used to input the target list data and the picking task to obtain the action block sequence and perform safety constraint filtering to drive the robotic arm to complete the picking action through a damage prevention control method, including: When the robotic arm moves to the target position according to the approach trajectory, the control mode is switched from position control to impedance control or force control mode; Specifically, in position control mode, the robotic arm moves along the planned trajectory to a predetermined distance (preferably 10-30mm) near the target fruit. During this stage, the end effector maintains free space movement and does not involve force feedback. When the tactile array or six-dimensional force sensor detects that the contact force exceeds a preset small threshold (such as 0.1-0.3N), the controller switches from position control mode to impedance control or force control mode. The switching process uses a gradual change in stiffness (S-shaped transition function) to avoid sudden impact.

[0045] The desired output force of the end effector is determined based on the deviation between the desired end effector pose and the current pose, as well as the feedforward force. Specifically, according to Calculate the desired output force at the end, where, The desired output force at the end; , These are the equivalent stiffness and damping matrix, respectively; , These are the desired end-effector pose and the current pose, respectively. , These are the expected terminal velocity and the current velocity, respectively. This is an estimate of the feedforward force (such as the preload generated by adsorption preload). The equivalent stiffness in the contact direction is set to a low value (preferably 50–200 N / m) to achieve compliant contact, while the stiffness in the tangential direction is appropriately increased (preferably 200–500 N / m) to prevent tangential slippage. The contact state is detected by a tactile sensor or a force sensor, and slippage is detected based on the ratio of the tangential force to the normal force of the desired output force at the end and the characteristics of tactile vibration energy. Specifically, the tactile array or force sensor provides the normal force. With tangential force and micro-vibration characteristics (Energy spectrum of high-frequency signals from the tactile array); calculated according to the slip discriminant function: ; in, The value of the slip discrimination function; For contact normal force; Contact tangential force; To prevent division by zero of small constants. By Vibrational energy characteristics Jointly determine the slippage trend.

[0046] Based on the slip detection results, adjust the clamping force or adsorption force in the desired output force at the end. After clamping is completed, perform the release action and place the target buffer into the collection box through the flexible fruit guide channel. Specifically, the controller determines whether the stable clamping conditions have been met based on the contact area and normal force fed back by the tactile array (such as the normal force reaching the minimum clamping force corresponding to the target fruit and the contact area exceeding the threshold). If the conditions are met, the current clamping force is maintained; if not, the clamping force is gradually increased (step by 0.1 to 0.5 N / time) until the conditions are met or the force limit protection is reached. After the clamping is stable, the release action is performed. The robotic arm exits the gap between the branches and leaves along the preset retraction trajectory, and then moves the end to the placement position. The fruit is placed into the collection box in a low-impact manner through the flexible fruit guide channel.

[0047] In practical applications, this embodiment utilizes impedance control and adaptive sliding adjustment. The end effector automatically switches to a compliant contact mode upon contact with the fruit, preventing damage to the peel caused by rigid impact. The sliding detection mechanism can sense the stability of fruit grasping in real time and dynamically adjust the clamping force to prevent fruit from falling off. The shearing or twisting combined with gentle pulling release method effectively reduces the impact torque transmitted to the fruit at the moment of stem breakage. Combined with the buffering release of the flexible fruit guide channel, a low-damage picking action sequence of "first stabilize, then release, and finally release slowly" is achieved, significantly reducing fruit compression damage and drop damage.

[0048] As an optional embodiment of the present invention, adjusting the clamping force or adsorption force in the desired output force at the end based on the slip detection result, and performing the release action after clamping is completed includes: When the ratio of the tangential force to the normal force or the tactile vibration energy characteristics exceeds the third preset threshold, it is determined that there is a slippage trend, and the clamping force or the adsorption force in the expected output force at the end is gradually increased until the slippage disappears or the upper limit is reached. Specifically, when the slip discrimination function value or vibrational energy characteristics When the third preset threshold is exceeded, a slippage trend is detected; the controller gradually increases the clamping force or adsorption force in increments of 0.1–0.5 N / step until the slippage signal disappears or the upper limit of force is reached; the upper limit of clamping force is dynamically determined based on the fruit diameter and maturity index. ; in, This is the upper limit of the dynamic clamping force; This is an estimate of the fruit diameter; It is a maturity index; , , The empirical coefficients were obtained by calibrating the fruits with different maturity levels and diameters through stress tests (the higher the maturity and the larger the diameter, the lower the upper limit of the dynamic clamping force should be to avoid excessive squeezing).

[0049] When the ratio of the tangential force to the normal force and the tactile vibration energy characteristics are both lower than the third preset threshold, the current clamping force or the adsorption force is maintained, and the stable clamping is determined to be completed. Specifically, when the slip discrimination function value or vibrational energy characteristics When all values ​​are below their respective preset thresholds, it is determined that the fruit has been stably clamped. The controller maintains the current clamping force or adsorption force unchanged and enters the detachment stage.

[0050] After stable clamping, the release action is performed by selecting either shearing and pulling or twisting and pulling, depending on the type of end effector.

[0051] Specifically, depending on the type of detachment mechanism configured in the end effector, if it is a shearing mechanism, the shearing blade cuts off the fruit stem and then pulls it gently at low speed in a predetermined direction to complete the separation; if it is a twisting mechanism, the twisting action breaks off the fruit stem and then pulls it gently to separate it. During the detachment process, the force control circuit continuously monitors the change in contact force to prevent the impact torque generated by the sudden breakage of the fruit stem from being transmitted to the fruit.

[0052] In practical applications, this embodiment utilizes a slip detection and adaptive clamping force adjustment mechanism to ensure that the end effector can stably grasp fruits of different sizes and ripeness with optimal clamping force, preventing fruit slippage due to insufficient clamping force or damage to the peel due to excessive clamping force. The dynamic force limit is adaptively adjusted according to fruit characteristics, with a lower force limit for more mature fruits, further reducing the risk of damage. The shearing or twisting combined with gentle pulling release method adapts to different fruit stalk characteristics, improving the success rate of harvesting and the rate of fruit integrity.

[0053] like Figure 6 and Figure 7 As shown, as an optional embodiment of the present invention, the training steps of the policy model include: Collect human demonstration data from multiple perspectives, including at least a first-person perspective, a third-person perspective, and a fixed perspective, and simultaneously record the joint angles of the robotic arm, the gripper status, and language annotations. Specifically, the demonstration data collection is as follows: human operators wear head-mounted cameras to collect first-person perspectives, handheld or mounted side-view cameras to collect third-person perspectives, and fixed perspectives to collect fixed perspectives through cameras fixed in the scene; during the operation, language annotations such as "prioritize picking ripe fruits" and "avoid branches" are input simultaneously through voice or button input; the joint angles and gripper states of the robotic arm are recorded in real time through the robot's own encoders and actuator states.

[0054] The human demonstration data is time-synchronized, the demonstration actions are mapped to the robot coordinate system, and training samples are constructed based on multimodal observations, language task vectors, and the action block sequence. The multimodal observations include at least the position of the target, the maturity index, and the uncertainty index. The language task vectors are obtained by encoding the language instructions for the picking task. Specifically, cross-view synchronization: using a certain video stream (preferably the end-camera view, as it is directly related to the robot's coordinate system) as the time reference, time offset estimation and interpolation synchronization are performed on the other video streams using feature point matching (ORB / SIFT features) or audio waveform alignment, so that multiple video streams correspond to the same physical event at the same time. State Reconstruction: For each synchronization moment, keyframe extraction (one frame every 0.1-0.2 s) combined with optical flow (such as RAFT) or feature matching is used to track the positional trajectories of the target fruit and the operator's hand / tool ​​in the image; visual odometry (such as ORB-SLAM3) is used to estimate the motion trajectory of the head-mounted / handheld camera relative to the scene, and then, through a fixed transformation between the pre-calibrated scene coordinate system and the robot coordinate system, the hand / tool ​​trajectory in the demonstration is mapped to the end effector pose trajectory in the robot coordinate system. Training sample construction: The synchronized multimodal observations (detection results, depth, semantic segmentation) are denoted as... Language / task annotations are encoded as vectors. The mapped end-effector pose increment, attitude increment, and gripping / adsorption / shearing commands are combined into an action block. to form training samples For data missing moments caused by occlusion or tracking loss, temporal interpolation (such as cubic spline interpolation or mask completion of temporal Transformer) is used to fill in the missing moments, and these moments are marked as "low confidence" samples, and their loss weight is reduced during training.

[0055] The policy model is trained based on the training samples, and the policy model is used to map multimodal observations and language task vectors into the action block sequence.

[0056] Specifically, the dataset can be expanded: in addition to successful demonstrations, failure samples (mistaken capture, slippage, collision, missed fragments) can be collected or labeled, and failure type labels can be recorded for comparative learning or weighted training to improve the strategy's ability to avoid failure modes. Policy model training: Input the constructed dataset into the selected policy model (Transformer sequence policy / VLA / diffusion model) and perform supervised training using behavioral cloning loss (such as mean squared error or negative log-likelihood).

[0057] The three strategy models are given explicit mathematical expressions, and the preferred implementation methods are explained.

[0058] Transformer sequence-to-sequence strategy: Set the historical observation window to the past. Step-by-step multimodal observation Language task vectors The encoder maps it to a context representation. The decoder outputs the future in an autoregressive or parallel manner. Step block The training objective is to minimize the mean square error of the demonstrated movement: ; in, The length of the historical observation window. To predict the step length of the movement.

[0059] Visual-Language-Action (VLA) Model: Basic policies are obtained through pre-training on large-scale general-purpose robotics datasets. ( For visual input, For language instructions, efficient parameter fine-tuning (such as LoRA) is used on blueberry picking data, updating only the low-rank adaptation matrix. ( The fine-tuning target is: ; in, For pre-trained model parameters, For LoRA fine-tuning the updated low-rank weight increments, For visual input, For language instructions, For blueberry picking dataset; Diffusion-based action generation model: Defines a forward noise addition process in the action space. ( , (Noise scheduling coefficients) are used to train the denoising network. Predict noise, For real action, To add noise The action following the step; The training loss is: ; Deployment via reverse denoising sampling Step by step, restore the movement .

[0060] The three models correspond to different engineering trade-offs: the Transformer model is lightweight and has fast inference speed, making it suitable for real-time deployment of edge computing units, and is recommended as the preferred implementation method in the independent claims; the VLA model relies on large-scale pre-training resources, has the strongest generalization ability but has a high inference latency; the diffusion model has the best action output diversity and is suitable for handling multimodal action distributions (such as multiple feasible approach angles for the same fruit), but the number of sampling steps leads to slower inference, which can be compressed to single digits through consistency distillation to meet real-time requirements.

[0061] In practical applications, the strategy model learns harvesting skills from human demonstrations, enabling it to adapt to fruit harvesting tasks under different varieties, greenhouse types, and lighting conditions. This reduces reliance on manual rules and parameter adjustments. At the same time, the safety constraint layer ensures the physical feasibility of the output actions, enhancing the system's generalization ability and engineering practicality in diverse orchard environments.

[0062] like Figure 6 and Figure 7 As shown, as an optional embodiment of the present invention, the low-damage harvesting method of multimodal sensing and cross-viewpoint mimicry further includes: The harvested data generated after deployment is written into the playback buffer. The harvested data includes at least successful cases and failed cases and their corresponding multimodal observations, action block sequences and result labels. Specifically, after the strategy model is deployed to the edge computing unit to perform the harvesting operation, the multimodal observations of each harvesting cycle, the action block sequence output by the strategy model, the execution result (success or failure), and the failure type label (such as mis-capture, slippage, collision, and missed harvest) are written into the replay buffer to form structured harvesting log data.

[0063] The data in the replay buffer is actively sampled based on the uncertainty index and failure type. Specifically, data samples with uncertainty indices higher than a preset threshold in the replay buffer, as well as data samples corresponding to various failure types, are selected first. After sorting the samples from high to low uncertainty, a predetermined proportion of samples are selected as high-value training samples, so that the training dataset is concentrated in scenarios where the model is currently performing poorly or has low perceived confidence.

[0064] By adjusting parameters, the strategy model is updated using the sampled data, forming a continuous learning loop with data feedback.

[0065] Specifically, a parameter-efficient fine-tuning method (such as LoRA) is adopted to periodically fine-tune and update the policy model with high-value training samples obtained from sampling. The updated model is then redeployed to the edge computing unit, forming a continuous learning closed loop of "demonstration learning - deployment verification - data feedback - model iteration".

[0066] In practical applications, this embodiment enables the harvesting robot to learn from its successes and failures during long-term operations through a continuous learning loop, gradually adapting to changes in the orchard environment (such as seasonal changes in sunlight, varietal differences, and changes in plant growth morphology). This allows for continuous optimization and adaptive updates of the model without human intervention, effectively improving the system's long-term operational stability and robustness.

[0067] like Figure 6 and Figure 7As shown, as an optional embodiment of the present invention, the low-damage harvesting method of multimodal sensing and cross-viewpoint mimicry further includes: Establish a parameterized orchard geometry and crop model, which includes at least row spacing, plant height, and branch and leaf density parameters, and obtain typical plant templates based on on-site scanning or structured light multi-view reconstruction. In a digital twin simulation environment, rendering data containing fruit distribution and maturity distribution is randomly generated. The rendering data includes at least RGB-D images, near-infrared images, and corresponding instance segmentation and three-dimensional position ground truth values. By perturbing the lighting, texture, noise, motion blur, occlusion, and rain / fog in the simulation environment through domain randomization, synthetic data is generated. The synthetic data is mixed with real-world collected data for training to expand the coverage of the training samples.

[0068] Specifically, a parameterizable crop model is established based on parameters such as orchard row spacing, plant height, and foliage density, and typical plant templates are obtained using on-site scanning or structured light multi-view reconstruction. Fruit distribution and maturity distribution are randomly generated in a twin environment to cover different harvesting seasons and management methods. RGB-D and near-infrared images are rendered in a simulation environment, and instance segmentation and 3D ground truth are directly output to avoid manual annotation costs. Domain randomization is used to perturb lighting, texture, noise, blur, occlusion, rain, fog, etc., to improve the model's generalization ability.

[0069] In practical applications, this embodiment generates large-scale synthetic data with ground truth annotations through a digital twin simulation environment, effectively solving the problems of high cost and limited coverage of real data collection and annotation. Domain randomization enables the model to pre-adapt to various lighting and occlusion conditions before deployment, reducing the on-site debugging cycle.

[0070] like Figure 6 and Figure 7 As shown, as an optional embodiment of the present invention, the low-damage harvesting method of multimodal sensing and cross-viewpoint mimicry further includes: Before deployment, the simulation model was style-aligned and noise statistically calibrated using a small set of real-world calibration images. The field data generated after deployment and the distribution of the real scene are fed back to the digital twin environment to update the parameterized orchard geometry and crop model and the domain randomized parameter distribution. The strategy model is efficiently fine-tuned by using a method of "pre-training through simulation followed by fine-tuning in real-world scenarios".

[0071] Specifically, before model deployment, a small set of real-world calibration data is used to align and calibrate the style, texture statistics, and noise distribution of the simulated data, making the synthesized data closer to the characteristics of real sensors. After the strategy model is deployed, the failed samples collected during execution and the new scene distribution data are sent back to the digital twin environment to correct the crop geometric parameters and domain randomization parameter distribution in the simulation model, so that the simulation environment gradually approximates the distribution of real operation scenarios. A training strategy of "pre-training in simulation first, then fine-tuning in reality" is adopted, in which the strategy model is pre-trained in the simulation environment first, and then deployed to a real orchard for efficient parameter fine-tuning.

[0072] In practical applications, Sim2Real calibration aligns simulation data with real sensor characteristics, reducing domain differences between simulation and real-world scenarios. The field feedback mechanism enables the digital twin environment to continuously track changes in the real-world scenario, thereby constantly correcting the distribution of simulation parameters and forming an engineering closed loop of "simulation enhancement - field feedback - iterative update".

[0073] like Figure 8 As shown, as an optional embodiment of the present invention, the low-damage harvesting method of multimodal perception and cross-viewpoint imitation further includes a multi-robot collaborative harvesting step: The plantation plots are divided into multiple operational zones according to row and column topology; The task allocation for each work zone is carried out through an auction-style scheduler in the cloud or edge. Each robot calculates the work cost to reach each zone based on its own travel distance, remaining power and expected revenue. The scheduler selects the robot-zone matching combination with the best overall cost to execute the task. Each robot independently performs harvesting operations within its assigned zone and reports its operational status to the scheduler in real time. By synchronizing the map and model parameters of each robot in real time through the edge station, all robots can share a consistent scene perception and decision-making capability. When the scheduler detects congestion or robot malfunction, it performs dynamic reallocation, adjusts the work zones or task order of each robot, and issues reallocation instructions to the corresponding robots. All robot operation data is aggregated at the edge station and uploaded to the cloud to generate cloud operation reports.

[0074] Specifically, in large-scale plantations, the system divides the plots into several work zones according to the row and column topology (row number, plant number) of the orchard. During auction-style allocation, each robot calculates its cost function to each zone. This cost includes at least the travel distance (the path length from the current position to the target zone), the remaining power (whether the current power is sufficient to complete the work in the zone), and the expected harvesting revenue in the zone (estimated based on the historical yield or real-time sensing results of the zone). The scheduler allocates each zone to the corresponding robot according to the principle of optimal comprehensive cost.

[0075] During operation, each robot synchronizes map and model parameters in real time through the edge station, reporting its own location, work progress, battery level, and abnormal status to the scheduler. The shared map information ensures that each robot avoids repeatedly covering the same area when navigating between rows, and the shared model parameters ensure that each robot's perception and decision-making criteria for the fruit are consistent.

[0076] If the scheduler detects that multiple robots are congested in a certain area (e.g., two robots enter the same work row at the same time) or that a robot malfunctions (e.g., insufficient power or mechanical failure), it will trigger dynamic reallocation: the unharvested targets in the congested area will be reassigned to nearby available robots, or the entire area that the malfunctioning robot has not completed will be reassigned to other robots, and a reallocation instruction will be issued.

[0077] All robot operation data is aggregated through the edge station and uploaded to the cloud to generate cloud operation reports containing information such as harvest yield (quantity and weight of fruit harvested in each zone), operation time (effective operation time and idle time of each robot), and fault records (fault type, occurrence time and handling result), providing data support for farm management.

[0078] In practical applications, this embodiment enables multiple robots to work collaboratively in large plantations through partitioned management and auction-based task allocation, avoiding redundant coverage and resource conflicts. The dynamic reallocation mechanism effectively addresses congestion and equipment failures during operation, improving the overall robustness and operational efficiency of the system. Edge stations aggregate maps and model parameters to ensure that all robots share consistent scene understanding and decision-making capabilities, while cloud-based reports provide data support for farm management.

[0079] like Figure 9 , Figure 10 , Figure 11 , Figure 12 and Figure 13As shown, exemplarily, the following describes a mechanical structure of a robot, which includes at least a mobile support system, a multi-joint manipulator system, and a modular end effector. The mobile support system includes a mobile body 3, a four-wheel omnidirectional independent chassis 4, and a lifting platform 1. The four-wheel omnidirectional independent chassis 4 uses Mecanum wheels or omnidirectional wheels to enable flexible movement between orchard rows and turning in place. The lifting platform 1 is installed in the middle of the body and is used to adjust the working height of the robot arm base according to the height of the plants. The mobile body 3 integrates an emergency stop switch 2, a touch screen 6, a lidar 5, and a storage frame 17. The multi-joint robot arm body 7 (the main component of the robot arm) is installed on the lifting platform 1 and has no less than 6 degrees of freedom. Its end is connected to the end effector 9 through a robot arm flange 10. Each joint of the robot arm body 7 integrates an encoder and a torque sensor 8 to provide joint angle and torque feedback.

[0080] The end effector 9 adopts a modular quick-change interface design and includes at least a flexible gripper 14, a flexible adsorption mechanism 12, and a flexible adsorption mechanism 2 13. The fingertips of the flexible gripper 14 are made of silicone or TPU soft material, and the gripping surface is set with micro-structure texture to increase friction and reduce indentation. The gripper opening range is 10-60mm, and the gripping force closed-loop range is 0-10N. The flexible adsorption mechanism 12 and the flexible adsorption mechanism 2 13 generate negative pressure through a micro vacuum pump or venturi device. The adsorption port is equipped with a flexible sealing ring, and the suction force closed-loop range is 0-6N, which is used to "point-suction stabilize" the fruit before gripping. The end effector 9 also integrates a fruit stem shearing or twisting component and a tactile or torque detection component to realize fruit stem separation and contact force sensing. A vision camera 11 is installed at the end of the robotic arm body 7 or at the top of the mast, including at least an RGB-D camera and a near-infrared camera, for acquiring multi-source harvesting observation data; the fruit guide groove 15 is a flexible fruit guiding channel, one end connected to the end effector 9, and the other end connected to the fruit storage basket 16. The inner wall of the channel is made of low-friction material and is provided with an energy-dissipating bending section, for cushioning the harvested fruit and placing it into the storage basket 16 in a low-impact manner; the storage basket 16 is placed in the storage frame 17, and the bottom of the storage frame 17 is provided with a storage basket guide 18 to facilitate the quick loading, unloading and positioning of the storage basket (for example, the storage basket guide 18 is at least one of a guide rail, guide groove or guide rod extending along the loading direction of the storage basket 16, and its inlet end is provided with a flared structure, for automatically correcting the loading angle when the storage basket 16 is pushed in, and locking the storage basket 16 by a limiting magnetic block after loading into place). The storage basket 16 can be provided with a vibration damping pad and a compartment structure to support the partitioning of the fruit according to maturity or size. All of the above components are communicatively connected to the edge computing and control unit, which executes the perception, decision-making, planning, and force control steps in the aforementioned method embodiments.

[0081] like Figure 14As shown, the present invention also provides a robot that applies the low-damage picking method of multimodal perception and cross-viewpoint imitation as described in the above embodiments, including: The perception module is used to: detect and segment multi-source harvesting observation data that has been time-synchronized and calibrated, output individual perception data that has been optimized by target enhancement and occlusion robustness, assess the maturity of targets and perform three-dimensional localization based on the individual perception data, and determine uncertainty indicators, so as to comprehensively output a target list data for harvesting. The cost assessment module is used to: comprehensively assess the harvesting benefits, collision risks, accessibility, detection uncertainty, and action switching costs based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, so as to obtain the comprehensive cost of each target; The collaborative planning module is used to: determine the current target based on the comprehensive cost, and based on the map accessibility data under the current target, to collaboratively plan the operation of the robotic arm and the movement of the chassis to determine the optimal stopping point and generate an approach trajectory; The picking execution module is used to: when the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, it uses a strategy model trained by cross-view imitation learning to input the target list data and picking task, obtains the action block sequence and performs safety constraint filtering, so as to drive the robotic arm to complete the picking action through the damage prevention control method.

[0082] The specific implementation method of this embodiment can be referred to the corresponding implementation method described above, and will not be described again here.

[0083] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0084] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

[0085] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A low-damage harvesting method using multimodal sensing and cross-viewpoint imitation, characterized in that, include: The multi-source harvesting observation data that has been time-synchronized and calibrated is detected and segmented, and individual perception data that has been optimized by target enhancement and occlusion robustness is output. Based on the individual perception data, the maturity of the target is assessed and three-dimensionally located, and uncertainty index is determined, so as to comprehensively output the target list data for harvesting. Based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, a comprehensive cost assessment is conducted on the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost to obtain the comprehensive cost of each target. The current objective is determined based on the comprehensive cost, and the operation of the robotic arm and the movement of the chassis are coordinated based on the map accessibility data under the current objective to determine the optimal stopping point and generate an approach trajectory. When the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, the strategy model trained by cross-view imitation learning is input with the target list data and the picking task to obtain the action block sequence and perform safety constraint filtering, so as to drive the robotic arm to complete the picking action through the damage prevention control method.

2. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 1, characterized in that, The multi-source harvesting and observation data includes at least RGB images, near-infrared images, and depth data; The process involves detecting and segmenting multi-source harvesting observation data that has undergone time synchronization and calibration, outputting individual perception data that has been optimized for target enhancement and occlusion robustness, assessing target maturity and performing 3D localization based on the individual perception data, and determining uncertainty indicators to comprehensively output a target list for harvesting, including: A multi-scale feature fusion network for densely occluded target scenes is used to detect and segment the RGB image, and output the target bounding box and instance segmentation mask; The near-infrared image of the detected fruit ROI region is fused with the RGB image, and a maturity index is output through a classifier or regressor. Based on the depth data, the pixel coordinates within the instance segmentation mask are transformed to three-dimensional coordinates in the robot base coordinate system to obtain the position of the target; The uncertainty index is determined based on at least one of detection confidence, depth noise, occlusion ratio, and reprojection error; The target list data is determined based on the location of each target, the maturity index, and the uncertainty index.

3. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 2, characterized in that, Also includes: When the uncertainty index exceeds the first preset threshold, the chassis is moved and / or the robotic arm is swung and scanned to update the position of the target and the maturity index; if the uncertainty index still exceeds the first preset threshold after the update, the target is downweighted during the sorting stage.

4. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 1, characterized in that, Based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, a comprehensive cost assessment is performed on the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost. The comprehensive cost of each target includes: The harvesting revenue is determined based on the maturity assessment results of each target; the collision risk is determined based on the relative positional relationship between each target and surrounding obstacles; the accessibility of each target is queried based on the accessibility map and the inverse accessibility map; and the action switching cost is determined based on the pose deviation between the current end pose and the poses of each target. The comprehensive cost of each objective is obtained by weighted summing of the harvesting revenue, collision risk, accessibility, detection uncertainty, and action switching cost.

5. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 1, characterized in that, The step of determining the current target based on the comprehensive cost, and coordinating the operation of the robotic arm and the movement of the chassis based on the map reachability data under the current target to determine the optimal stopping point and generate an approach trajectory includes: If the current docking point is reachable from the target and the overall risk is not higher than the second preset threshold, the robotic arm trajectory planning is directly triggered to generate the approach trajectory from the current end pose to the target picking pose; If the current docking point is unreachable from the target or the overall risk is higher than the second preset threshold, a set of candidate docking points is searched on the inverse reachability map. Each docking point in the candidate docking point set is comprehensively evaluated based on the number of reachable targets, average movement cost, driving distance, and turning difficulty indicators, and the optimal docking point is selected by weighting. The chassis is driven to move to the optimal docking point, and local fine positioning is performed by visual markers or ground structures after the chassis is in place. After positioning is complete, the robotic arm trajectory planning is triggered to generate the approach trajectory from the current end pose to the target picking pose.

6. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 5, characterized in that, When the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, the strategy model trained through cross-view imitation learning is used to input the target list data and the picking task to obtain the action block sequence and perform safety constraint filtering to drive the robotic arm to complete the picking action through damage prevention control. When the robotic arm moves to the target position according to the approach trajectory, the control mode is switched from position control to impedance control or force control mode; The desired output force of the end effector is determined based on the deviation between the desired end effector pose and the current pose, as well as the feedforward force. The contact state is detected by a tactile sensor or a force sensor, and slippage is detected based on the ratio of the tangential force to the normal force of the desired output force at the end and the characteristics of tactile vibration energy. Adjust the clamping force or adsorption force in the desired output force at the end according to the slip detection results, perform the release action after clamping is completed, and release the target buffer into the collection box through the flexible fruit guide channel.

7. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 6, characterized in that, The step of adjusting the clamping force or adsorption force in the desired output force at the end based on the slip detection result, and performing the release action after clamping, includes: When the ratio of the tangential force to the normal force or the tactile vibration energy characteristics exceeds the third preset threshold, it is determined that there is a slippage trend, and the clamping force or the adsorption force in the expected output force at the end is gradually increased until the slippage disappears or the upper limit is reached. When the ratio of the tangential force to the normal force and the tactile vibration energy characteristics are both lower than the third preset threshold, the current clamping force or the adsorption force is maintained, and the stable clamping is determined to be completed. After stable clamping, the release action is performed by selecting either shearing and pulling or twisting and pulling, depending on the type of end effector.

8. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 2, characterized in that, The training steps for the policy model include: Collect human demonstration data from multiple perspectives, including at least a first-person perspective, a third-person perspective, and a fixed perspective, and simultaneously record the joint angles of the robotic arm, the gripper status, and language annotations. The human demonstration data is time-synchronized, the demonstration actions are mapped to the robot coordinate system, and training samples are constructed based on multimodal observations, language task vectors, and the action block sequence. The multimodal observations include at least the position of the target, the maturity index, and the uncertainty index. The language task vectors are obtained by encoding the language instructions for the picking task. The policy model is trained based on the training samples, and the policy model is used to map multimodal observations and language task vectors into the action block sequence.

9. The low-damage harvesting method for multimodal sensing and cross-viewpoint imitation according to claim 8, characterized in that, Also includes: The harvested data generated after deployment is written into the playback buffer. The harvested data includes at least successful cases and failed cases and their corresponding multimodal observations, action block sequences and result labels. The data in the replay buffer is actively sampled based on the uncertainty index and failure type. By adjusting parameters, the strategy model is updated using the sampled data, forming a continuous learning loop with data feedback.

10. A robot, characterized in that, The low-damage harvesting method using multimodal sensing and cross-viewpoint mimicry as described in any one of claims 1-9 includes: The perception module is used to: detect and segment multi-source harvesting observation data that has been time-synchronized and calibrated, output individual perception data that has been optimized by target enhancement and occlusion robustness, assess the maturity of targets and perform three-dimensional localization based on the individual perception data, and determine uncertainty indicators, so as to comprehensively output a target list data for harvesting. The cost assessment module is used to: comprehensively assess the harvesting benefits, collision risks, accessibility, detection uncertainty, and action switching costs based on the target list data, the robot's reachable docking area data, and the robotic arm's reachable posture data, so as to obtain the comprehensive cost of each target; The collaborative planning module is used to: determine the current target based on the comprehensive cost, and based on the map accessibility data under the current target, to collaboratively plan the operation of the robotic arm and the movement of the chassis to determine the optimal stopping point and generate an approach trajectory; The picking execution module is used to: when the robot executes the cooperative planning to reach the optimal stopping point and drives the robotic arm to move to the target position according to the approach trajectory, it uses a strategy model trained by cross-view imitation learning to input the target list data and picking task, obtains the action block sequence and performs safety constraint filtering, so as to drive the robotic arm to complete the picking action through the damage prevention control method.