A vision system and method for a cone picking platform

By employing a dual-stage visual collaboration scheme and a lidar-assisted drone harvesting system, high-precision pine cone target identification and maturity assessment were achieved in complex forest environments. This solved the problems of low identification accuracy and insufficient positioning precision in existing systems, thereby improving the efficiency and safety of harvesting operations.

CN122228840APending Publication Date: 2026-06-19HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2026-03-20
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing drone harvesting systems suffer from low target recognition accuracy, difficulty in determining maturity, and insufficient harvesting positioning precision in complex forest environments, resulting in low harvesting efficiency and high fruit damage rate, making it difficult to achieve efficient and precise pine cone harvesting.

Method used

A two-stage visual collaboration scheme is adopted, combining a wide-angle gimbal camera and a hand-eye camera. It uses a detection algorithm guided by prior information to identify long-distance targets, and achieves near-range maturity judgment and accurate positioning through multi-dimensional feature fusion. It integrates LiDAR for environmental perception and path planning.

Benefits of technology

It significantly improved the accuracy of pine cone identification and maturity assessment, reduced the fruit damage rate, increased the automation and environmental adaptability of harvesting operations, and enhanced operational safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122228840A_ABST
    Figure CN122228840A_ABST
Patent Text Reader

Abstract

A vision system and method for a pine cone harvesting platform, belonging to the field of intelligent forestry equipment, addresses the problems of low target recognition accuracy, difficulty in maturity assessment, and insufficient harvesting positioning precision in existing drone harvesting systems operating in complex forest environments. The system includes a vision processing unit, a wide-angle gimbal camera, a LiDAR, and a hand-eye camera. The wide-angle gimbal camera acquires distant visual images. The vision processing unit uses a priori information-guided detection method to identify pine cone targets and combines LiDAR point cloud data to plan the platform's motion path. The hand-eye camera, mounted at the end of a robotic arm, acquires near-field images and depth information after the platform reaches the target area. The vision processing unit uses an attention-weighted mechanism to fuse color, size, and scale flipping angle features to determine maturity, locate the strike point, and control the harvesting actuator to complete the operation. This invention achieves high-precision identification and positioning of pine cone targets through two-stage visual collaboration and multi-dimensional feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a vision system and its usage method for a pine cone harvesting platform, belonging to the fields of forestry intelligent equipment, machine vision and unmanned aerial vehicle autonomous operation technology. Background Technology

[0002] Pine cones, an important forestry economic resource, are rich in pine nuts and have high edible and medicinal value, widely used in food processing, health products, and the forestry economic industry chain. Pine cones are mostly distributed in the middle and upper layers of the tree canopy, and the harvesting process faces technical challenges such as high-altitude operations, limited working space, small target size, and severe shading from branches and leaves. Traditional harvesting methods struggle to balance efficiency, safety, and fruit quality.

[0003] Currently, pine cone harvesting mainly relies on the following methods:

[0004] Manual tree climbing: Workers climb trees and use long poles or pick the fruit by hand. This method is characterized by high risk, high labor intensity, and low efficiency. Moreover, with the decrease in rural labor force, labor costs continue to rise, making it difficult to meet the needs of large-scale harvesting.

[0005] High-pole knocking method: Using a long pole to strike pine cones to make them fall is suitable for the lower canopy area, but it has limited effect on harvesting pine cones in the middle and upper layers, and the cones are easily damaged during the striking process, affecting the commercial value of pine nuts.

[0006] Vibratory harvesting machinery: Although it can improve harvesting efficiency, it causes non-selective impact on the tree and immature fruit, affecting the pine tree's continuous production capacity.

[0007] Lifting platform auxiliary method: The lifting equipment is used to send the workers to the height of the tree canopy for picking. Although it reduces the risk of climbing trees to a certain extent, the equipment is expensive, has poor mobility, is difficult to deploy flexibly in forest environments with large slopes and complex terrain, and has high requirements for the access conditions of the work area.

[0008] In recent years, with the development of drone technology and intelligent sensing technology, some studies have attempted to use multi-rotor drones equipped with harvesting devices for pine cone harvesting. However, existing drone harvesting systems still face the following key technical bottlenecks in complex forest environments:

[0009] Low target recognition accuracy: Factors such as variable lighting conditions in the forest, cluttered backgrounds, and severe obstruction by branches and leaves make it difficult for the visual perception system to reliably identify pine cone targets;

[0010] The difficulty in determining maturity: existing systems cannot effectively distinguish between mature and immature pine cones, which can easily lead to the misharvesting of immature cones and affect the sustainable use of forestry resources;

[0011] Insufficient harvesting positioning accuracy: The lack of a precise spatial positioning mechanism makes it difficult for the harvesting execution agency to accurately target the pine cones.

[0012] The aforementioned problems make it difficult for existing drone harvesting systems to achieve efficient and accurate pine cone harvesting in actual forest operations, thus hindering the promotion and application of intelligent forestry equipment.

[0013] Therefore, there is an urgent need to develop a vision system that can achieve high-precision target identification, accurate maturity judgment, and precise harvesting positioning of pine cones in complex forest environments, so as to improve the accuracy and reliability of harvesting operations, reduce the fruit damage rate, and protect the sustainable production capacity of forestry resources. Summary of the Invention

[0014] To address the problems of low target recognition accuracy, difficulty in determining maturity, and insufficient picking positioning precision in existing drone harvesting systems in complex forest environments, this invention provides a vision system and method for pine cone harvesting platforms.

[0015] In a first aspect, the present invention provides a vision system for a pine cone harvesting platform, the platform comprising a main platform and a harvesting unit and a vision system disposed thereon, the harvesting unit comprising a robotic arm and a harvesting execution mechanism; the vision system comprising a vision processing unit 1, a wide-angle gimbal camera 2, a lidar 4 and a hand-eye camera 3;

[0016] The vision processing unit 1 is used to receive and process the sensor data returned by the wide-angle gimbal camera 2, the lidar 4 and the hand-eye camera 3, and output control commands to the work platform, the robotic arm and the picking execution mechanism.

[0017] The wide-angle gimbal camera 2 is mounted on the main platform via a gimbal structure and is communicatively connected to the vision processing unit 1. The wide-angle gimbal camera 2 is used to adjust its attitude under the control of the vision processing unit 1, acquire far-end visual images of the working area, and transmit the far-end visual images to the vision processing unit 1. The far-end visual images are used for long-distance identification and preliminary positioning of the pine cone target.

[0018] The lidar 4 is located in the upper part of the main platform and is communicatively connected to the vision processing unit 1. The lidar 4 is used to scan the surrounding environment under the scheduling and control of the vision processing unit 1, generate three-dimensional point cloud data and transmit it to the vision processing unit 1. The three-dimensional point cloud data is used to construct a working environment model and provide environmental perception information for the path planning and obstacle avoidance of the working platform.

[0019] The vision processing unit 1 identifies the target location of the pine cone based on the remote vision image, and plans the motion path in conjunction with the working environment model to generate control commands to move the working platform to the target working area.

[0020] The hand-eye camera 3 is installed at the end of the robotic arm and is spatially associated with the picking mechanism to form a hand-eye structure. The hand-eye camera 3 is connected to the vision processing unit 1. The hand-eye camera 3 is used to acquire near-end image information and depth information of the working area at the end of the robotic arm after the working platform arrives at the target working area and the robotic arm adjusts its posture, and transmits the near-end image information and depth information to the vision processing unit 1. The near-end image information and depth information are used for near-range maturity identification of pine cone targets and precise positioning of the striking location.

[0021] The vision processing unit 1 determines the relative pose of the robotic arm end relative to the target pine cone and the pine cone maturity based on the near-end image information and depth information, and controls the robotic arm to drive the picking execution mechanism to the target working pose to perform the picking operation.

[0022] Preferably, the operating platform is a coaxial octocopter flight platform or a forestry operation robot;

[0023] The coaxial octocopter flight platform includes a frame system, a power battery pack, an electronic speed controller, a brushless motor device, a propeller assembly, a flight controller, an attitude sensor unit, and a GPS positioning sensor, with the frame system serving as the main platform; the forestry operation robot includes a mobile chassis, a power system, a control system, and navigation sensors, with the mobile chassis serving as the main platform.

[0024] The main platform of the work platform is used to carry the vision processing unit 1, the wide-angle gimbal camera 2, the hand-eye camera 3 and the lidar 4, and to realize the movement and attitude control of the work platform.

[0025] Preferably, the vision processing unit 1 includes a processor and a memory, used to receive sensor data collected by the wide-angle gimbal camera 2, the hand-eye camera 3 and the lidar 4, and to perform fusion processing on the sensor data to generate target positioning information, maturity assessment information, strike position information and platform motion planning information, and to output control commands to the work platform, the robotic arm and the harvesting execution mechanism.

[0026] The vision processing unit 1 interacts with the work platform, robotic arm and picking mechanism via CAN bus, communicates with the wide-angle gimbal camera 2, hand-eye camera 3 and lidar 4 via serial communication interface, and communicates remotely with the external control terminal via wireless communication module.

[0027] Preferably, the lidar 4 is a multi-line lidar scanning radar, used to scan the surrounding environment under the scheduling and control of the vision processing unit 1, generate three-dimensional point cloud data and transmit it to the vision processing unit 1; the three-dimensional point cloud data is used to construct a working environment model, providing environmental perception information for the path planning, obstacle avoidance and spatial positioning of the target working area of ​​the working platform.

[0028] Preferably, the wide-angle gimbal camera 2 is mounted on the main platform via a controllable gimbal structure and is configured to perform attitude adjustment and long-distance visual image acquisition under the control of the visual processing unit 1, and the visual images are transmitted to the visual processing unit 1.

[0029] The visual processing unit 1 is configured to perform long-range target recognition, and uses a prior information-guided detection method to identify pine cone targets; the prior information-guided detection method includes:

[0030] A multidimensional Gaussian prior distribution model of pine cone morphological features is pre-constructed, wherein the morphological features include spherical shape parameters, scale texture features and size range;

[0031] The prior distribution model is input into the target detection network and modulated with the depth feature map extracted by the detection network to generate a modulated feature map.

[0032] Candidate regions are generated based on the modulated feature map, and the posterior probability of each candidate region relative to the prior distribution is calculated.

[0033] Based on the posterior probability, candidate regions are filtered and non-maximum suppression is applied to output the detection box coordinates and size information of the pine cone target.

[0034] Preferably, the hand-eye camera 3 is installed at the end of the robotic arm and is spatially associated with the picking execution mechanism to form a hand-eye structure; the hand-eye camera 3 is used to acquire near-end image information and depth information of the working area at the end of the robotic arm after the working platform arrives at the target working area and the robotic arm adjusts its posture, and transmits the near-end image information and depth information to the vision processing unit 1.

[0035] The vision processing unit 1 is configured to extract features from the near-end image information, including multi-dimensional features such as color, size, and scale flipping angle; input the multi-dimensional features into an attention weighting network, and perform weighted fusion of the multi-dimensional features through learnable attention weights to obtain a fused feature vector; input the fused feature vector into a maturity classifier to output the pine cone maturity level; for pine cones determined to be mature, the coordinates of the impact point in the image coordinate system are obtained through regression by a key point detection network, and combined with the depth information to be transformed into the coordinate system of the robotic arm end effector to generate an impact position command.

[0036] Preferably, the coordinate systems of each functional module establish a spatial reference relationship through the coordinate system of the robotic arm end effector to realize the spatial position association between the hand-eye camera 3 and the picking execution mechanism; the vision processing unit 1 is configured to perform a unified transformation on the coordinate systems of the wide-angle gimbal camera 2, the hand-eye camera 3, the lidar 4, the work platform, and the robotic arm base to realize the spatial alignment of multi-source sensor data.

[0037] In a second aspect, the present invention provides a harvesting method using a vision system for a pine cone harvesting platform, comprising the following steps:

[0038] Step 1: The wide-angle gimbal camera 2 adjusts its posture under the control of the vision processing unit 1, acquires far-end visual images of the work area, and transmits the far-end visual images to the vision processing unit 1.

[0039] Step 2: Based on the received remote visual image, the visual processing unit 1 uses a detection algorithm guided by prior information to identify the pinecone target, obtain the rough location information of the pinecone target, and combine it with the environmental point cloud data obtained by the lidar 4 to plan the target hovering position and movement path of the work platform.

[0040] Step 3: After the work platform moves to the target hovering position, the robotic arm adjusts its end effector posture so that the hand-eye camera 3 is aligned with the area where the target pine cone is located; the hand-eye camera 3 acquires near-end image information and depth information of the working area at the end of the robotic arm and transmits it to the vision processing unit 1;

[0041] Step 4: The visual processing unit 1 identifies the maturity of the pine cone based on the received near-end image information and depth information, extracts color features, size features and scale flipping angle features, performs multi-dimensional feature fusion, determines the maturity level of the pine cone, and determines the specific striking location.

[0042] Step 5: The vision processing unit 1 generates control commands based on the determined striking position and sends them to the work platform and robotic arm via the CAN bus to drive the picking mechanism to complete the picking operation; after the operation is completed, the posture is adjusted to enter the next work cycle.

[0043] Preferably, in step two, the detection algorithm guided by prior information includes pre-constructing a multidimensional Gaussian prior distribution model of pine cone morphological features, inputting the prior distribution model into the target detection network for feature modulation, generating candidate regions based on the modulated feature map and calculating the posterior probability, and filtering candidate regions according to the posterior probability to output the detection box coordinates and size information of the pine cone target.

[0044] Preferably, in step four, maturity identification includes inputting the extracted multidimensional features into an attention weighting network for weighted fusion, inputting the fused feature vector into a maturity classifier to output the pine cone maturity level, and for pine cones determined to be mature, obtaining the strike point coordinates through key point detection network regression, and combining depth information to convert to the coordinate system of the robotic arm end effector.

[0045] The beneficial effects of this invention are:

[0046] (1) High recognition accuracy: This system adopts a dual-stage vision scheme that combines a wide-angle gimbal camera and a hand-eye camera. At long distances, prior information is used to guide detection to overcome occlusion and background interference. At close distances, multi-dimensional feature fusion is used to ensure the accuracy of target locking, which significantly improves the pine cone recognition rate in complex forest environments.

[0047] (2) Accurate maturity judgment: This system introduces an attention weighting mechanism through hand-eye cameras, and integrates color, size and morphological characteristics to accurately identify the maturity of pine cones, avoid the mis-harvesting of immature fruits and protect the sustainable production capacity of forestry resources.

[0048] (3) High positioning accuracy: This system obtains the depth information of the target working area through a depth camera and establishes a spatial reference relationship by combining the coordinate system of the end of the robotic arm, thereby achieving accurate positioning of the pine cone striking position and reducing the fruit damage rate.

[0049] (4) Strong environmental adaptability: This system integrates the environmental perception function of lidar and combines it with the vision system to realize obstacle avoidance and path planning in the forest, which enhances the safety of the system in complex unstructured environments.

[0050] (5) High degree of automation: This system realizes full-process automated control from target search, maturity assessment to picking location, reducing reliance on human experience and improving picking efficiency. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the structure of a vision system for a pine cone harvesting platform according to the present invention;

[0052] Figure 2 for Figure 1 Side view;

[0053] Figure 3 This is a flowchart illustrating the overall workflow of a vision system for a pine cone harvesting platform according to the present invention.

[0054] Figure 4 A flowchart illustrating the process of target positioning and platform planning;

[0055] Figure 5 A flowchart illustrating the process of maturity identification and target positioning.

[0056] Explanation of markings in the diagram:

[0057] 1. Visual processing unit; 2. Wide-angle gimbal camera; 3. Hand-eye camera; 4. LiDAR. Detailed Implementation

[0058] This invention addresses the problems of low target recognition accuracy, difficulty in determining maturity, and insufficient harvesting positioning precision in existing drone-based pine cone harvesting systems in complex forest environments. It proposes a pine cone harvesting vision system based on two-stage visual collaboration. The core innovation of this invention lies in:

[0059] 1. Dual-stage collaborative detection architecture: A two-level visual perception system of "far-near" is formed by a wide-angle gimbal camera and a hand-eye camera to achieve progressive recognition from large-scale target search to local precise positioning;

[0060] 2. Prior information-guided recognition algorithm: In the long-distance recognition stage, prior information on pine cone morphology features is introduced. The target features are described by probability distribution, which improves the recognition accuracy in occluded or low-discrimination environments.

[0061] 3. Attention-weighted maturity assessment: In the close-range recognition stage, an attention-weighted mechanism is adopted, which integrates multi-dimensional features such as color, size, and scale flipping angle to achieve accurate identification of pine cone maturity and precise determination of the strike position.

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] I. System Overall Structure

[0064] Reference Figure 1 and Figure 2 As shown, the present invention proposes a vision system for a pine cone harvesting platform, comprising a vision processing unit 1, a wide-angle gimbal camera 2, a lidar 4, and a hand-eye camera 3. The platform includes a main platform and a harvesting unit mounted thereon, the harvesting unit comprising a robotic arm and a harvesting execution mechanism.

[0065] The vision processing unit 1 is mounted on the main platform and serves as the core computing carrier of the system. It receives and processes sensor data returned by the wide-angle gimbal camera 2, LiDAR 4, and hand-eye camera 3, and outputs control commands to the work platform, robotic arm, and harvesting mechanism. The vision processing unit 1 employs an embedded high-performance computing module with edge computing capabilities, used for real-time processing of multi-source sensor data and generation of control commands. The vision processing unit 1 includes a processor and memory. The processor runs deep learning models and various algorithms, while the memory stores program code and temporary data.

[0066] The wide-angle gimbal camera 2 is mounted on the main platform via a gimbal structure and is communicatively connected to the vision processing unit 1. The wide-angle gimbal camera 2 employs a large field-of-view lens with a horizontal field of view of no less than 90°, enabling it to cover a large area of ​​the tree canopy in a single shot, thus improving target search efficiency. The gimbal structure has two degrees of freedom for adjustment: pitch and yaw. Under the control of the vision processing unit 1, it can adjust its attitude to achieve active search and tracking of targets in different directions. The wide-angle gimbal camera 2 is used for attitude adjustment under the control of the vision processing unit 1, acquiring remote visual images of the work area, and transmitting these remote visual images to the vision processing unit 1. These remote visual images are used for long-range identification and preliminary localization of pine cone targets.

[0067] The lidar 4 is located on the upper part of the main platform and is communicatively connected to the vision processing unit 1. Lidar 4 is a multi-line laser scanning radar used to scan the surrounding environment under the scheduling and control of the vision processing unit 1, generating 3D point cloud data and transmitting it to the vision processing unit 1. The 3D point cloud data is used to construct a working environment model, providing environmental perception information for the platform's path planning and obstacle avoidance. During actual operation, lidar 4 continuously scans the environment, and the generated point cloud data, after denoising, filtering, and downsampling, is used to construct a 3D map of the working environment, enabling the platform to autonomously locate itself and detect obstacles in unknown environments.

[0068] The vision processing unit 1 identifies the target location of the pine cone based on the remote vision image, and plans the motion path in conjunction with the working environment model to generate control commands to move the working platform to the target working area.

[0069] The hand-eye camera 3 is installed at the end of the robotic arm and is spatially associated with the harvesting mechanism to form a hand-eye structure. This means the camera's field of view covers the working area of ​​the harvesting mechanism, enabling "eye-to-hand" visual servo control. The hand-eye camera 3 is connected to the vision processing unit 1 and is used to acquire near-end image and depth information of the working area at the end of the robotic arm after the work platform reaches the target working area and the robotic arm adjusts its posture. This near-end image and depth information is then transmitted to the vision processing unit 1. The near-end image and depth information are used for close-range maturity identification of pine cones and precise positioning of the striking location. The hand-eye camera 3 can simultaneously acquire color images and depth information, providing a data foundation for subsequent image processing and target localization.

[0070] The vision processing unit 1 determines the relative pose of the robotic arm end relative to the target pine cone and the pine cone maturity based on the near-end image information and depth information, and controls the robotic arm to drive the picking execution mechanism to the target working pose to perform the picking operation.

[0071] The vision processing unit 1 interacts with the work platform, robotic arm and picking mechanism via CAN bus, communicates with the wide-angle gimbal camera 2, hand-eye camera 3 and lidar 4 via serial communication interface, and communicates with the external control terminal remotely via wireless communication module to upload work status and receive task instructions.

[0072] The coordinate systems of each functional module establish a spatial reference relationship through the coordinate system of the robotic arm's end effector, realizing the spatial position association between the hand-eye camera 3 and the picking execution mechanism. The vision processing unit 1 is configured to perform unified transformations on the coordinate systems of the wide-angle gimbal camera 2, the hand-eye camera 3, the lidar 4, the work platform, and the robotic arm base, achieving spatial alignment of multi-source sensor data. Specifically, the system obtains the transformation matrix between each coordinate system through calibration. When the hand-eye camera 3 detects the target pine cone, it transforms the target's pose in the camera coordinate system to the robotic arm's end effector coordinate system, then to the robotic arm base coordinate system, and finally to the world coordinate system, achieving spatial positioning of the target.

[0073] II. Operating Platform Adaptation

[0074] The operating platform can be a coaxial octocopter flight platform or a forestry operation robot, depending on the actual operating scenario.

[0075] When using a coaxial octocopter flight platform, its hardware components include a frame system, a power battery pack, an electronic speed controller, a brushless motor unit, propeller assemblies, a flight controller, an attitude sensor unit, and a GPS positioning sensor, with the frame system serving as the main platform. The coaxial octocopter layout employs a two-layer coaxial arrangement of rotors, with four rotors in each layer, for a total of eight rotors. This provides greater lift within the same size, while also exhibiting good wind resistance and redundancy, making it particularly suitable for stable hovering operations in complex forest environments.

[0076] When forestry robots are used, their hardware components include a mobile chassis, a power system, a control system, and navigation sensors, with the mobile chassis serving as the main platform. Forestry robots are suitable for ground-based, walking operations in forests and can harvest pine cones in gently sloping woodlands.

[0077] The main platform of the work platform is used to carry the vision processing unit 1, the wide-angle gimbal camera 2, the hand-eye camera 3 and the lidar 4, and to realize the movement and attitude control of the work platform.

[0078] III. Long-range target identification and platform planning

[0079] Reference Figure 4 As shown, the detailed process of long-range target recognition and platform planning in this invention is as follows:

[0080] Step S201: Remote Image Acquisition

[0081] Under the control of the vision processing unit 1, the wide-angle gimbal camera 2 adjusts its posture, acquires far-field visual images of the work area, and transmits these images to the vision processing unit 1. The vision processing unit 1 controls the gimbal movement according to a preset scanning strategy to ensure full coverage of the target tree canopy area.

[0082] Step S202: Prior Information-Guided Target Detection

[0083] The visual processing unit 1 processes the received far-end visual image and uses a detection method guided by prior information to identify the pine cone target. Specifically, it includes the following sub-steps:

[0084] First, a multidimensional Gaussian prior distribution model of pine cone morphological characteristics is pre-constructed. These morphological characteristics include spherical shape parameters (e.g., aspect ratio close to 1:1), scale texture characteristics (regularity of scale arrangement), and size range (typical pine cone diameter range, such as 5-15 cm). This prior information is obtained through statistical analysis of a large number of pine cone samples, forming a multidimensional Gaussian distribution model that describes the probability distribution of pine cone characteristics.

[0085] Secondly, the prior distribution model is input into the target detection network and modulated with the depth feature map extracted by the detection network. Specifically, the detection network performs convolution operations on the input image to extract multi-scale depth feature maps; the prior distribution model is converted into a modulation factor, which is then multiplied element-wise with the feature map, making the network focus more on regions conforming to the prior distribution. This feature modulation mechanism is equivalent to introducing prior knowledge of pinecone morphology into the network, guiding the network to enhance its response to potential pinecone regions during feature extraction.

[0086] Then, candidate regions are generated based on the modulated feature map, and the posterior probability of each candidate region relative to the prior distribution is calculated. The detection network generates several candidate regions on the modulated feature map through a region proposal network, each candidate region corresponding to a location in the image where a pine cone may be present. For each candidate region, the network extracts its corresponding feature vector and calculates the confidence score of the region belonging to the pine cone category through a fully connected layer. Simultaneously, combined with the prior distribution model, the posterior probability of the candidate region relative to the prior distribution is calculated using Bayes' theorem, that is, the probability that the region is a pine cone after comprehensively considering the confidence score output by the network and the prior information.

[0087] Finally, candidate regions are filtered and non-maximum suppression is applied based on the posterior probability, and the detection box coordinates and size information of the pinecone targets are output. Specifically, a posterior probability threshold is set, and candidate regions higher than the threshold are retained; for overlapping candidate regions, a non-maximum suppression algorithm is used to retain the region with the highest posterior probability, eliminating duplicate detections. The detection box coordinates (such as center point coordinates and bounding box size) of each pinecone target are finally output for subsequent use.

[0088] Step S203: Platform Status Awareness

[0089] The vision processing unit 1 acquires the current position, attitude, and battery status data of the work platform. This data comes from the platform's onboard positioning sensors (such as GPS), attitude sensors (IMU), and battery monitoring module, and is used to assess the platform's current mobility and remaining working time.

[0090] Step S204: Hovering Path Planning

[0091] The vision processing unit 1 combines the pine cone target location information output in step S202 with the platform status data obtained in step S203 to calculate the optimal hovering point and motion path. The selection principle for the hovering point is to ensure that the robotic arm can cover the target pine cone with minimal joint movement after deployment, while ensuring sufficient safety distance between the working platform and the tree. Path planning adopts... The algorithm, or RRT algorithm, combined with the environmental point cloud data provided by LiDAR 4, generates a collision-free motion path from the current position to the target hovering point.

[0092] Step S205: Platform Attitude Adjustment

[0093] The vision processing unit 1 generates control commands based on the planned motion path and sends them to the platform controller via the CAN bus. This controls the work platform to move to the planned hovering position and adjust its attitude to stabilize and align with the target area. After reaching the hovering point, the platform maintains stable hovering through closed-loop control of position and attitude, creating conditions for subsequent robotic arm operations.

[0094] IV. Close-range maturity identification and strike positioning

[0095] Reference Figure 5 As shown, the detailed process of close-range maturity identification and strike positioning in this invention is as follows:

[0096] Step S301: Acquisition of near-end image and depth information

[0097] After the working platform reaches the target hovering position, the robotic arm adjusts its end effector posture under the control of the vision processing unit 1, so that the hand-eye camera 3 is aligned with the area where the target pine cone is located. The hand-eye camera 3 acquires a near-end color image and depth information of the working area of ​​the robotic arm end effector, and transmits the near-end image information and depth information to the vision processing unit 1.

[0098] Step S302: Multidimensional Feature Extraction

[0099] The visual processing unit 1 performs feature extraction on the received near-end image information, extracting multi-dimensional features including color features, size features, and scale flip angle features.

[0100] Color feature extraction: The image is converted to the HSV color space, and the distribution histograms of hue, saturation, and brightness of the pine cone region are statistically analyzed to form a color feature vector. Mature pine cones are usually yellowish-brown or brownish-yellow, while immature pine cones are mostly bluish-green. Color features play an important role in indicating maturity.

[0101] Size feature extraction: Based on the target detection bounding box, the pixel area, aspect ratio, and equivalent diameter of the pine cone are calculated. Size features reflect the growth degree of the pine cone; mature pine cones typically reach a certain size range.

[0102] Scale flipping angle feature extraction: This is a key innovative feature of this invention. The scales of mature pine cones flip outwards, forming a unique morphological feature. By calculating the rate of change of convexity of the pine cone outline, or by using an edge detection algorithm to extract the gradient direction of the scale edges, the angular distribution of the scales relative to the pine cone surface is statistically analyzed. Mature pine cones have larger scale flipping angles, while the scales of immature pine cones are close to the pine cone surface, with an angle close to 0°. This feature is a key indicator for judging maturity.

[0103] Step S303: Attention-weighted fusion

[0104] The multidimensional features extracted in step S302 are input into an attention-weighted network, which then performs weighted fusion of the multidimensional features using learnable attention weights. The attention-weighted network employs a lightweight neural network structure, containing fully connected layers and a softmax activation function. The network takes the multidimensional feature vector as input and outputs attention weights for each feature dimension. These weights are automatically learned during training via backpropagation. The weights reflect the importance of each feature to maturity assessment under the current environmental conditions: for example, color features may receive higher weights under good lighting conditions; and scale flipping angle features may receive higher weights when foliage occlusion causes color distortion. The original features are then weighted and summed with the attention weights to obtain the fused feature vector.

[0105] Step S304: Output of maturity level

[0106] The fused feature vector is input into a maturity classifier, which outputs the pine cone maturity level. The maturity classifier uses a multilayer perceptron structure, with the output layer employing a softmax activation function to output the probabilities of two categories: mature and immature. The classifier determines whether the pine cone has reached the harvestable maturity level based on the fused feature vector, avoiding the misharvesting of immature pine cones.

[0107] Step S305: Target location

[0108] For pine cones determined to be mature in step S304, visual processing unit 1 obtains the coordinates of the impact point in the image coordinate system through keypoint detection network regression. The keypoint detection network adopts an encoder-decoder structure, using color images captured by a hand-eye camera as input, and outputs a heatmap representing the probability distribution of the impact point. The network is specially trained to accurately locate the optimal impact position on the pine cone—usually selecting the center point of the pine cone or the connection point of the stalk, as these locations are most likely to cause the pine cone to fall off with minimal damage after impact.

[0109] Step S306: Coordinate Transformation and Command Generation

[0110] The vision processing unit 1 transforms the image coordinates of the impact point obtained in step S305, combined with the depth information acquired by the hand-eye camera 3, into the coordinate system of the robotic arm's end effector. The specific transformation process is as follows: based on the intrinsic parameter matrix of the hand-eye camera 3, the image coordinates and depth values ​​are converted into three-dimensional point coordinates in the camera coordinate system; based on the fixed transformation matrix between the hand-eye camera 3 and the robotic arm's end effector (obtained through hand-eye calibration), the coordinates in the camera coordinate system are transformed into the robotic arm's end effector coordinate system; then, based on the kinematic model of the robotic arm, the target point in the end effector coordinate system is transformed into the robotic arm's base coordinate system. Finally, motion control commands for each joint of the robotic arm are generated, enabling the picking actuator to move precisely to the impact point position.

[0111] Step S307: Harvesting

[0112] The vision processing unit 1 sends the generated control commands to the robotic arm controller via the CAN bus, driving the picking actuator to move to the target working position and perform the picking operation. The picking actuator can be a clamping, shearing, or impact type, selecting the appropriate picking method according to the specific situation of the pine cone. After picking is completed, the hand-eye camera 3 acquires the post-operation image for confirmation. If the pine cone is successfully detached, the system prepares to enter the next work cycle; if it fails, it can choose to re-identify or abandon the target according to the preset strategy.

[0113] V. Overall System Workflow

[0114] Reference Figure 3 As shown, the overall workflow of the vision system of this invention integrates the two stages of long-distance recognition and close-range identification, forming a complete operational loop:

[0115] Step S101: Long-distance target search

[0116] Under the control of the vision processing unit 1, the wide-angle gimbal camera 2 adjusts its posture, acquires far-field visual images of the work area, and transmits these images to the vision processing unit 1. This stage utilizes the wide field of view of the wide-angle gimbal camera to quickly scan a large area of ​​the tree canopy and detect potential pine cone targets.

[0117] Step S102: Target Positioning and Platform Planning

[0118] Based on the received far-end visual image, the vision processing unit 1 uses a detection algorithm guided by prior information to identify the pinecone target, obtains the rough location information of the pinecone target, and combines it with the environmental point cloud data obtained by the lidar 4 to plan the target hovering position and movement path of the work platform. For the specific implementation of this stage, please refer to Part Three (…). Figure 4 ) detailed description.

[0119] Step S103: Platform movement and robotic arm adjustment

[0120] After the work platform moves to the target hovering position, the robotic arm adjusts its end effector to align the hand-eye camera 3 with the area where the target pine cone is located. The platform remains stable after reaching the hovering point, creating conditions for subsequent precise operations.

[0121] Step S104: Maturity Identification and Targeting

[0122] The hand-eye camera 3 acquires near-end image information and depth information of the working area at the end of the robotic arm and transmits it to the vision processing unit 1. Based on the received near-end image information and depth information, the vision processing unit 1 identifies the maturity of the pine cone, extracts color features, size features, and scale flipping angle features, performs multi-dimensional feature fusion, determines the pine cone maturity level, and identifies the specific striking location. For the specific implementation of this stage, please refer to Part Four (…). Figure 5 ) detailed description.

[0123] Step S105: Harvesting Execution and Cycle

[0124] The vision processing unit 1 generates control commands based on the determined impact position and sends them to the work platform controller and the robotic arm controller via the CAN bus, driving the harvesting actuator to complete the harvesting operation. After harvesting is completed, the system determines whether there are any unprocessed pine cone targets: if so, it returns to step S101 to enter the next work cycle; if all are completed or the battery is insufficient, the work platform automatically returns to its starting position.

[0125] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.

Claims

1. A vision system for a pine cone harvesting platform, characterized in that, The operating platform includes a main platform and a picking unit and a vision system mounted thereon. The picking unit includes a robotic arm and a picking execution mechanism. The vision system includes a vision processing unit (1), a wide-angle gimbal camera (2), a lidar (4), and a hand-eye camera (3). The vision processing unit (1) is used to receive and process the sensor data returned by the wide-angle gimbal camera (2), lidar (4) and hand-eye camera (3), and output control commands to the work platform, robotic arm and picking execution mechanism. The wide-angle gimbal camera (2) is mounted on the main platform through a gimbal structure and is connected to the vision processing unit (1) in communication. The wide-angle gimbal camera (2) is used to adjust its posture under the control of the vision processing unit (1), collect remote visual images of the working area, and transmit the remote visual images to the vision processing unit (1). The remote visual images are used for long-distance identification and preliminary positioning of the pine cone target. The lidar (4) is located in the upper part of the main platform and is connected to the vision processing unit (1). The lidar (4) is used to scan the surrounding environment under the scheduling and control of the vision processing unit (1), generate three-dimensional point cloud data and transmit it to the vision processing unit (1). The three-dimensional point cloud data is used to construct the working environment model and provide environmental perception information for the path planning and obstacle avoidance of the working platform. The visual processing unit (1) identifies the target location of the pine cone based on the remote visual image, and plans the motion path in conjunction with the working environment model, and generates control commands to move the working platform to the target working area. The hand-eye camera (3) is installed at the end of the robotic arm and is spatially associated with the picking mechanism to form a hand-eye structure. The hand-eye camera (3) is connected to the vision processing unit (1). The hand-eye camera (3) is used to acquire near-end image information and depth information of the working area at the end of the robotic arm after the working platform arrives at the target working area and the robotic arm adjusts its posture, and transmits the near-end image information and depth information to the vision processing unit (1). The near-end image information and depth information are used for near-range maturity identification of pine cone targets and precise positioning of the striking position. The vision processing unit (1) determines the relative pose of the end of the robotic arm with respect to the target pine cone and the maturity of the pine cone based on the near-end image information and depth information, and controls the robotic arm to drive the picking execution mechanism to move to the target working pose to perform the picking operation.

2. The vision system for a pine cone harvesting platform according to claim 1, characterized in that, The operating platform is a coaxial octagonal flight platform or a forestry operation robot; The coaxial octocopter flight platform includes a frame system, a power battery pack, an electronic speed controller, a brushless motor device, a propeller assembly, a flight controller, an attitude sensor unit, and a GPS positioning sensor, with the frame system serving as the main platform; the forestry operation robot includes a mobile chassis, a power system, a control system, and navigation sensors, with the mobile chassis serving as the main platform. The main platform of the work platform is used to carry the vision processing unit (1), the wide-angle gimbal camera (2), the hand-eye camera (3) and the lidar (4), and to realize the movement and attitude control of the work platform.

3. The vision system for a pine cone harvesting platform according to claim 2, characterized in that, The vision processing unit (1) includes a processor and a memory, used to receive sensor data collected by the wide-angle gimbal camera (2), hand-eye camera (3) and lidar (4), and to perform fusion processing on the sensor data to generate target positioning information, maturity assessment information, strike position information and platform motion planning information, and output control commands to the work platform, robotic arm and picking execution mechanism. The vision processing unit (1) interacts with the work platform, robotic arm and picking execution mechanism via CAN bus, communicates with the wide-angle gimbal camera (2), hand-eye camera (3) and lidar (4) via serial communication interface, and communicates with the external control terminal remotely via wireless communication module.

4. The vision system for a pine cone harvesting platform according to claim 1, characterized in that, The lidar (4) is a multi-line laser scanning lidar, used to scan the surrounding environment under the scheduling and control of the vision processing unit (1), generate three-dimensional point cloud data and transmit it to the vision processing unit (1); the three-dimensional point cloud data is used to construct the working environment model, and provide environmental perception information for the path planning, obstacle avoidance and spatial positioning of the target working area of ​​the working platform.

5. A vision system for a pine cone harvesting platform according to claim 1, characterized in that, The wide-angle gimbal camera (2) is mounted on the main platform through a controllable gimbal structure and is configured to perform attitude adjustment and long-distance visual image acquisition under the control of the visual processing unit (1), and the visual image is transmitted to the visual processing unit (1). The visual processing unit (1) is configured to perform long-distance target recognition and to identify pine cone targets using a detection method guided by prior information; The detection method based on prior information includes: A multidimensional Gaussian prior distribution model of pine cone morphological features is pre-constructed, wherein the morphological features include spherical shape parameters, scale texture features and size range; The prior distribution model is input into the target detection network and modulated with the depth feature map extracted by the detection network to generate a modulated feature map. Candidate regions are generated based on the modulated feature map, and the posterior probability of each candidate region relative to the prior distribution is calculated. Based on the posterior probability, candidate regions are filtered and non-maximum suppression is applied to output the detection box coordinates and size information of the pine cone target.

6. The vision system for a pine cone harvesting platform according to claim 1, characterized in that, The hand-eye camera (3) is installed at the end of the robotic arm and is spatially associated with the picking execution mechanism to form a hand-eye structure. The hand-eye camera (3) is used to acquire near-end image information and depth information of the working area at the end of the robotic arm after the working platform arrives at the target working area and the robotic arm adjusts its posture, and transmits the near-end image information and depth information to the vision processing unit (1). The visual processing unit (1) is configured to extract features from the near-end image information, including multi-dimensional features such as color features, size features and scale flip angle features; The multidimensional features are input into an attention-weighted network, and the multidimensional features are weighted and fused through learnable attention weights to obtain a fused feature vector. The fused feature vector is input into a maturity classifier to output the pine cone maturity level. For pine cones that are determined to be mature, the coordinates of the impact point in the image coordinate system are obtained by regression through a key point detection network, and the coordinates are transformed into the coordinate system of the robotic arm end effector by combining the depth information to generate an impact position command.

7. A vision system for a pine cone harvesting platform according to claim 1, characterized in that, The coordinate systems of each functional module establish a spatial reference relationship through the coordinate system of the end of the robotic arm, thereby realizing the spatial position association between the hand-eye camera (3) and the picking execution mechanism; the vision processing unit (1) is configured to perform a unified transformation on the coordinate system of the wide-angle gimbal camera (2), the coordinate system of the hand-eye camera (3), the coordinate system of the lidar (4), the coordinate system of the work platform and the coordinate system of the robotic arm base, thereby realizing the spatial alignment of multi-source sensor data.

8. A harvesting method using a vision system for a pine cone harvesting platform, characterized in that, The vision system according to any one of claims 1 to 7 includes the following steps: Step 1: The wide-angle gimbal camera (2) adjusts its posture under the control of the vision processing unit (1), acquires the far-end visual image of the working area, and transmits the far-end visual image to the vision processing unit (1). Step 2, Visual processing unit (1) Based on the received remote visual image, it uses a detection algorithm guided by prior information to identify the pinecone target, obtains the rough position information of the pinecone target, and combines the environmental point cloud data obtained by the lidar (4) to plan the target hovering position and movement path of the work platform. Step 3: After the work platform moves to the target hovering position, the robotic arm adjusts the end posture so that the hand-eye camera (3) is aligned with the area where the target pine cone is located; the hand-eye camera (3) acquires the near-end image information and depth information of the working area at the end of the robotic arm and transmits it to the vision processing unit (1). Step 4, Visual Processing Unit (1) Based on the received near-end image information and depth information, the maturity of the pine cone is identified, color features, size features and scale flipping angle features are extracted and multi-dimensional feature fusion is performed to determine the maturity level of the pine cone and the specific striking location is determined. Step 5: The vision processing unit (1) generates control commands based on the determined striking position and sends them to the work platform and robotic arm via the CAN bus to drive the picking execution mechanism to complete the picking operation; after the operation is completed, it adjusts its posture to enter the next work cycle.

9. A harvesting method using a vision system for a pine cone harvesting platform according to claim 8, characterized in that, In step two, the detection algorithm guided by prior information includes pre-constructing a multidimensional Gaussian prior distribution model of pine cone morphological features, inputting the prior distribution model into the target detection network for feature modulation, generating candidate regions based on the modulated feature map and calculating the posterior probability, and filtering candidate regions according to the posterior probability to output the detection box coordinates and size information of the pine cone target.

10. A harvesting method using a vision system for a pine cone harvesting platform according to claim 8, characterized in that, In step four, maturity identification includes inputting the extracted multidimensional features into an attention weighting network for weighted fusion, inputting the fused feature vector into a maturity classifier to output the pine cone maturity level, and for pine cones determined to be mature, regressing the strike point coordinates through a key point detection network and combining the depth information to convert them to the coordinate system of the robotic arm end effector.