Target fruit dynamic tracking method and system based on visual perception and deep learning

Through the collaborative operation of large-field-of-view and small-field-of-view sensors and deep learning methods, the problem of inaccurate target detection of fruit picking robots in complex orchard environments was solved, and high-precision fruit positioning and picking decisions were achieved.

CN120747736AActive Publication Date: 2025-10-03AGRI MACHINERY INST CHINESE TROPICAL ACAD OF SCI +2

Patent Information

Application Number
CN202510830918.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing fruit picking robot systems find it difficult to simultaneously take into account macroscopic perception and microscopic precise positioning in complex orchard environments, especially in situations such as lighting changes, shadow interference, and fruit occlusion, which leads to inaccurate target detection and affects the accuracy of picking decisions.

Method used

A method based on visual perception and deep learning is adopted. Through the collaborative operation of large-field-of-view sensors and small-field-of-view sensors, an adaptive control model of the encoder-decoder architecture and a multi-target segmentation network are combined to achieve multi-scale, multi-angle, and multi-resolution collaborative perception. Pseudo point cloud data and BEV views are constructed using RGB-D information. Dynamic target tracking is performed in combination with the Kalman filter algorithm, and a picking suitability scoring function is constructed.

Benefits of technology

It significantly improves the fruit detection coverage and positioning accuracy, can adapt to changing lighting conditions, accurately identify semi-obscured fruits, solves the problem of inaccurate dynamic target tracking, and realizes scientific picking point selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747736A_ABST
    Figure CN120747736A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target tracking, and discloses a target fruit dynamic tracking method and system based on visual perception and deep learning. The method comprises the following steps: collecting original image data, carrying out image enhancement through an encoder-decoder architecture adaptive regulation and control model, carrying out deep learning target detection on an enhanced large-vision-field image to extract fruit position information, guiding a small-vision-field sensor to a picking preparation point, and carrying out deep learning target detection on the fruit position information; the method comprises the following steps: acquiring fruit shielding rate and growth attitude information through a multi-target segmentation network, performing dynamic tracking in combination with depth information to obtain motion trail prediction data, calculating a picking suitability score and an optimal picking point according to the shielding rate, the growth attitude and the motion trail, and generating a picking execution instruction. The fruit detection coverage rate and the positioning precision in a complex orchard environment are improved, and the limitation of single picking point selection in the prior art is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target tracking technology, and in particular to a method and system for dynamic tracking of target fruits based on visual perception and deep learning. Background Art

[0002] Robotic automatic recognition and harvesting technology has garnered widespread attention in the fruit harvesting sector. However, the complexity of orchard environments presents significant challenges for automated harvesting. Existing fruit harvesting robotic systems generally rely on single vision sensors and traditional image processing algorithms. While this design can achieve basic object recognition under ideal conditions, it has significant limitations in practical orchard applications.

[0003] Orchard environments face challenges such as fluctuating lighting, shadows, fruit obscured by branches and leaves, and wind-induced fruit swaying, making it difficult for a single vision system to simultaneously achieve both macroscopic perception and microscopic precision positioning. Traditional image processing methods are particularly prone to overexposure or underexposure in strong or low light conditions, resulting in loss of target features. Furthermore, when fruit is partially obscured, existing object detection algorithms often struggle to accurately identify its complete form and growth posture, compromising the accuracy of picking decisions. Summary of the Invention

[0004] The present invention provides a method and system for dynamic tracking of target fruits based on visual perception and deep learning. The present invention improves the fruit detection coverage and positioning accuracy in complex orchard environments, and overcomes the limitation of single picking point selection in the existing technology.

[0005] In a first aspect, the present invention provides a method for dynamic tracking of a target fruit based on visual perception and deep learning, the method comprising:

[0006] Collect the original image data and perform image enhancement through the adaptive control model of the encoder-decoder architecture to obtain the enhanced large-viewing-area image;

[0007] Performing deep learning target detection on the enhanced large-viewing-area image, extracting the location information of the fruit target and the location of the picking preparation point, and guiding the preset small-viewing-area sensor to the location of the picking preparation point;

[0008] Collect local image data and perform pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit target;

[0009] Dynamically tracking the fruit target according to the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data;

[0010] According to the occlusion rate, the growth posture information and the motion trajectory prediction data, a picking suitability score and an optimal picking point are calculated, and a picking execution instruction is generated.

[0011] In a second aspect, the present invention provides a target fruit dynamic tracking system based on visual perception and deep learning, the target fruit dynamic tracking system based on visual perception and deep learning comprises:

[0012] The image enhancement module is used to collect raw image data and perform image enhancement through an adaptive control model of the encoder-decoder architecture to obtain an enhanced large-viewing-area image;

[0013] A target detection module is used to perform deep learning target detection on the enhanced large-viewing-area image, extract the location information of the fruit target and the location of the picking preparation point, and guide the preset small-viewing-area sensor to the location of the picking preparation point;

[0014] The pixel-level segmentation module is used to collect local image data and perform pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit target;

[0015] A dynamic tracking module is used to dynamically track the fruit target based on the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data;

[0016] A calculation module is used to calculate the picking suitability score and the optimal picking point according to the occlusion rate, the growth posture information and the motion trajectory prediction data, and generate a picking execution instruction.

[0017] The technical solution provided by this invention overcomes the technical bottleneck of a single visual sensor's inability to simultaneously achieve global perception and local precision identification through the collaborative operation of a large-field-of-view sensor and a small-field-of-view sensor. This solution achieves multi-scale, multi-angle, and multi-resolution collaborative perception, significantly improving fruit detection coverage and positioning accuracy in complex orchard environments. An adaptive image control model employing an encoder-decoder architecture effectively addresses the impact of interference factors such as strong light, shadows, occlusion, and motion blur on image quality in orchard environments. Two dedicated processing branches perform pixel-level enhancement and global ISP parameter learning, respectively, enabling the system to adapt to varying lighting conditions. A deep learning-based object detection and multi-target segmentation network enhances the ability to extract fruit features. A multi-target instance segmentation network with a MASNet architecture effectively identifies partially occluded fruit, providing accurate occlusion rate and fruit posture information for subsequent picking decisions. Combining RGB-D information to construct pseudo point cloud data and BEV views enables fruit positioning in three-dimensional space. A Kalman filter algorithm incorporating an adaptive process noise adjustment mechanism accurately predicts the motion trajectory of fruit under wind disturbances, addressing the inaccurate tracking of dynamic targets by traditional methods. By comprehensively considering multiple key factors such as fruit occlusion rate, growth posture, depth value, motion state and complexity of the surrounding environment, a scientific picking suitability scoring function was constructed; the working parameters of the end effector were automatically adjusted according to different picking difficulties, overcoming the limitation of single picking point selection in existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 Schematic diagram of an embodiment of a method for dynamic tracking of a target fruit based on visual perception and deep learning in an embodiment of the present invention;

[0020] Figure 2 This is a schematic diagram of an embodiment of a target fruit dynamic tracking system based on visual perception and deep learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] An embodiment of the present invention provides a method and system for dynamic tracking of target fruits based on visual perception and deep learning. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices.

[0022] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 In one embodiment of the present invention, a method for dynamic tracking of a target fruit based on visual perception and deep learning includes:

[0023] Step S101: collecting original image data and performing image enhancement through an adaptive control model of an encoder-decoder architecture to obtain an enhanced large-viewing-area image;

[0024] Specifically, a large-field-of-view sensor in a cascaded vision system is used to capture raw images of the orchard environment. The sensor continuously acquires image data of a large-scale scene within a set wide-angle field of view at a resolution of 1920×1080 and a frequency of 30 frames per second. The raw image data is grid-divided, dividing the entire image into several overlapping sub-region image blocks of a fixed size (e.g., 160×160 pixels), with the overlap rate between sub-regions set to 20%. This operation effectively addresses the problem of local perceptual distortion caused by local illumination intensity changes, shadow interference, and uneven imaging in the image. The resulting set of divided image sub-regions is sequentially input into an image adaptive control model with an encoder-decoder structure. Based on the perception of image quality changes in different regions, this model can implement a parallel and coordinated processing strategy of local enhancement and global unified adjustment. The encoder layer of the model uses a five-layer continuous convolutional architecture to extract multi-level features from subregions of the input image. Each convolution layer uses a 3×3 kernel with a stride of 2 and a number of channels, followed by 64, 128, 256, 512, and 512. Batch normalization and activation functions are used to ensure the stability and nonlinear representation of feature extraction. The encoding process compresses the original image layer by layer into a low-dimensional feature space representation, extracting discriminative high-level semantic features from the image. After feature encoding, the encoded features are fed into the decoder for restoration. The decoder also consists of five layers, but uses transposed convolutions to gradually upsample the feature maps and reconstruct the image. Each transposed convolution layer maps the encoded low-dimensional features to a high-resolution space, restoring the dimensional structure of the original image. Skip connections are introduced in the encoding layer during the decoding process to preserve low-level detail information, resulting in higher structural restoration and texture fidelity in the initial decoded result. The initial decoding results are fed into two independent processing branches for parallel processing. The first branch focuses on pixel-level image enhancement. Three pixel enhancement modules are connected in series. Each module adopts a residual structure and introduces a dilated convolution operation to expand the receptive field. This enhances the representation of blurred edges, weak contrast, and occluded areas while maintaining the integrity of image details. The final output is a set of pixel-level enhanced features. Meanwhile, the second branch learns image signal processing parameters. Using an attention mechanism, it acquires the global illumination and color information of the input image and generates a global adjustment parameter set consisting of a color matrix and gamma correction values. This parameter set is used to express the overall control requirements for the entire image in terms of color distribution, brightness balance, and tone curve. The pixel-level enhancement features output by the first processing branch are fused with the global adjustment parameters output by the second processing branch. The fusion operation includes weighted superposition and back-mapping. This ensures that the image retains the clear details after local enhancement while maintaining the overall consistency of color and illumination characteristics, resulting in an enhanced large-viewing-area image.

[0025] Step S102: Perform deep learning target detection on the enhanced large-viewing-area image to extract the location information of the fruit target and the location of the picking preparation point, and guide the preset small-viewing-area sensor to the location of the picking preparation point;

[0026] Specifically, the enhanced large-viewing-area image is fed into an object detection network for feature extraction. This network, designed with a lightweight architecture, employs depthwise separable convolutions as its backbone. By decomposing the standard convolution operation into two stages: channel-wise and point-wise convolutions, this reduces the network's computational complexity and parameter count while maintaining sensitivity to the image's spatial structure and channel features. This architecture processes the enhanced image to produce an initial feature map that retains the semantic information of the fruit. This initial feature map is then fed into a CBAM module with an integrated attention mechanism for joint modeling of channel and spatial attention. In the channel attention mechanism, the network performs average and max pooling operations on the feature map along the spatial dimension and maps it through a fully connected layer to form a channel weight vector, thereby highlighting key channel features relevant to fruit recognition. In the spatial attention mechanism, the feature map is pooled along the channel dimension and combined with convolution operations to produce a spatial response weight map, emphasizing edge contours, boundaries, and high-response regions within the image. By cascading these two attention mechanisms, the resulting enhanced feature map exhibits enhanced representational capabilities and localization capabilities. The enhanced feature map is fed into three different-scale detection heads for object detection. Each head identifies target features within a specific scale range. A multi-scale fusion detection mechanism ensures consistent performance when detecting small, occluded, distant, and close-up fruits. Each head outputs a set of information, including fruit class probabilities, bounding box coordinates, and confidence scores. By aggregating the multi-scale detection results, the positional distribution of the target fruit on the image plane is obtained. To eliminate overlapping frames and redundant recognition, non-maximum suppression is performed on the detection results. The bounding boxes with the highest confidence scores and low overlap with other candidate boxes are selected as the final output, generating a set of fruit target bounding boxes with accurate boundaries and no redundant overlap. Based on the fruit target bounding box information, the coordinates of the center point of each fruit bounding box are calculated. This center point serves as the pre-harvest point for the fruit. The corresponding depth value is extracted from the depth channel of the wide-viewing-area sensor. This 2D image center point is combined with the spatial depth information to construct a 3D representation of the pre-harvest point. After acquiring this spatial position, a series of coordinate transformation operations are performed to map this point from the perception level to the control level. The pre-picking preparation point is transformed from the image space to the robot's body reference coordinate system via a pre-established rigid body transformation matrix between the large-field-of-view sensor and the robot's reference coordinate system. This is then further transformed into the picking mechanism's motion control coordinate system via a transformation matrix defined by the robot's structural parameters, completing the coordinate chain mapping from the perception point to the execution point and generating a set of pose control parameters for precisely driving the small-field-of-view sensor. Based on these pose control parameters, the robot's motion control unit guides the small-field-of-view sensor to quickly and accurately move to the designated pre-picking preparation point through interpolation calculations and real-time feedback, ensuring that the target fruit area is included in its field of view.

[0027] Step S103: collecting local image data and performing pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit object;

[0028] Specifically, a small-view sensor is guided to the target area using the pre-picking points provided by the large-view vision module. Within this viewing angle, high-resolution local image data is collected. This image contains complex details such as the fruit body, occluding branches and leaves, the fruit stem, and adjacent background textures. This image is then fed into a multi-object segmentation network for initial processing. Based on a pixel-level recognition mechanism, the network outputs an initial pixel-level mask. This mask semantically delineates the fruit, leaves, and background at the image level, marking the initial boundary between the possible fruit region and the interfering regions. To improve the accuracy of boundary determination and the hierarchical robustness of segmentation, the ResNext-101 framework, with its strong feature extraction capabilities, is used as the backbone architecture. A feature pyramid network is then combined for multi-scale feature extraction, effectively representing all types of objects in the image at different spatial scales. This generates a feature atlas with high resolution and multi-level contextual information. Based on this multi-scale feature atlas, bounding box regression and classification are performed. The network regression module determines the bounding box position of each candidate object, and the classification module verifies its fruit classification, resulting in a set of accurately identified bounding box information. Based on the area defined by the bounding box, the system reactivates the mask generation branch of the segmentation network. Using local feature up- and down-sampling, skip connections, and convolutional fusion, the system constructs a refined mask. This mask accurately restores the complete structure of the fruit at the pixel level and effectively removes interference caused by edge blur and partial occlusion, resulting in a precise segmentation mask for the fruit. The occlusion rate of the fruit object is calculated based on the refined segmentation mask. By statistically analyzing the difference between the actual visible area in the mask and the overall predicted area, the system identifies whether the fruit is significantly occluded from the given viewing angle. The higher the occlusion rate, the more difficult it is to fully identify and safely harvest the fruit. Simultaneously, the system extracts edge contour information from the refined mask and constructs a set of boundary points. Directional modeling is performed based on the image coordinate structure. The angle between the principal direction vector and the image vertical is calculated to determine the angle between the fruit's current growth posture and the direction of gravity. This angle information indicates the fruit's tilt, growth orientation, and suitability for harvesting from the current angle. These two structural pieces of information, occlusion rate and growth posture, are integrated into the system's harvesting decision module to contribute to fruit harvestability scoring and prioritization.

[0029] Step S104: Dynamically track the fruit target based on the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data;

[0030] Specifically, a small-field-of-view sensor is used to capture image data containing a depth channel, generating an RGB-D image. Each pixel in the RGB-D image carries both color information and a corresponding depth value. A coordinate backprojection operation is performed on each pixel in the RGB-D image based on the camera's intrinsic parameter matrix, converting its 2D image coordinates into 3D spatial coordinates. This creates a pseudo-point cloud data structure containing both spatial position and color information. Based on the pseudo-point cloud data of the fruit region, a horizontal XOY coordinate system is established with the picking mechanism as the origin. This coordinate system, projected onto the horizontal plane, maps the fruit's position in the 3D coordinate system onto a 2D bird's-eye-view plane, forming a BEV (bird's-eye-view) coordinate representation. This representation facilitates motion trajectory calculation and path fitting during dynamic tracking. Simultaneously, the coordinates of the fruit's bounding box center point and the corresponding depth value are extracted from the segmentation results from the small-field-of-view sensor. By analyzing the bounding box center point within the segmentation mask and combining it with the depth information at that point, a 3D coordinate representation of the fruit within the local field of view is obtained. This data is then correlated with the constructed BEV spatial position to determine the 3D spatial position of the fruit target. The three-dimensional spatial position of the fruit target is matched between frames. By comparing the differences in the fruit's spatial position between consecutive frames, the fruit's trajectory path during motion is identified and a complete motion trajectory is constructed by combining time series data. A Kalman filter algorithm is used to dynamically estimate the trajectory state. This algorithm uses the fruit's spatial position and velocity state as internal variables and predicts and updates the state between each frame based on temporal evolution. By combining the observed and predicted states, a smoother and more continuous fruit trajectory curve is generated, resulting in fruit state prediction data. An adaptive process noise adjustment mechanism is introduced based on the fruit state prediction data. By monitoring the amplitude of the fruit's position change between adjacent image frames, the process noise covariance parameter in the filter is dynamically adjusted. Specifically, when the fruit moves violently and has large displacements between consecutive frames, the system's response to uncertainty is improved, increasing the error tolerance of the prediction region. When the fruit moves slowly or tends to be stable, the covariance value is reduced to enhance prediction accuracy. Based on this state estimation and dynamic adjustment model, the fruit trajectory prediction data is output.

[0031] Step S105: Calculate the picking suitability score and the optimal picking point based on the occlusion rate, growth posture information, and motion trajectory prediction data, and generate a picking execution instruction.

[0032] Specifically, the multi-source fruit status information output by the preceding modules is comprehensively processed and integrated into a structured set of fruit picking factors. This information includes the fruit occlusion rate calculated from pixel-level segmentation results, the growth attitude angle obtained from edge morphology analysis, the fruit depth value and motion state output by the trajectory tracking module, and the surrounding environment complexity index calculated by the environmental perception unit. These factors collectively describe the fruit's pickability, picking difficulty, and picking risk under the current environment and time. Based on this factor set, a multidimensional picking suitability scoring function is constructed. The scoring function converts each factor into a standardized index value through weighting or model mapping, and then comprehensively calculates these factors to generate a picking suitability score. This score reflects whether the target fruit has a high picking priority under the current conditions. All identified fruits are ranked by suitability score, and a scoring threshold is set to select a set of fruits suitable for immediate picking as the list of fruits to be picked. For each fruit to be picked, a fine structure analysis process is initiated. Based on a high-precision segmentation mask from a small field of view image, the fruit's edge features and internal contours are extracted. The stem connection area is then identified. The stem connection point is located in the narrow area where the upper or lateral part of the fruit meets the branch. This point is considered a key reference point for the picking gripper to contact the target. After identifying the stem connection point, a local coordinate system is established on the fruit surface. Based on the spatial relationship between this point and the fruit's geometric center, the optimal approach path and optimal picking point position are calculated, enabling the picking device to achieve stable gripping with the shortest path and minimal interference angle. Based on the optimal picking point and the dynamic state of the fruit, as reflected by the picking suitability score, the execution strategy is customized. For example, for target fruits with high scores, a high-speed, low-damping mode is assigned for rapid picking. For fruits with borderline scores, a medium-speed, steady-state control strategy is adopted to balance safety and efficiency. The spatial position, clamping posture, approach direction and control mode parameters of the optimal picking point are encapsulated into a standardized picking execution instruction, which is sent to the end effector through the control link, and the motion trajectory and clamping strategy are adjusted in real time according to the feedback during the picking process.

[0033] In an embodiment of the present invention, the collaborative operation of a large-field-of-view sensor and a small-field-of-view sensor overcomes the technical bottleneck of a single visual sensor's inability to simultaneously achieve global perception and local precision identification. This achieves multi-scale, multi-angle, and multi-resolution collaborative perception, significantly improving fruit detection coverage and positioning accuracy in complex orchard environments. An adaptive image control model employing an encoder-decoder architecture effectively addresses the impact of interference factors such as strong light, shadows, occlusion, and motion blur on image quality in orchard environments. Two dedicated processing branches perform pixel-level enhancement and global ISP parameter learning, respectively, enabling the system to adapt to changing lighting conditions. A deep learning-based target detection and multi-target segmentation network enhances the ability to extract fruit features. The multi-target instance segmentation network of the MASNet architecture effectively identifies semi-occluded fruit, providing accurate occlusion rate and fruit posture information for subsequent picking decisions. Combining RGB-D information to construct pseudo point cloud data and BEV views enables fruit positioning in three-dimensional space. A Kalman filter algorithm incorporating an adaptive process noise adjustment mechanism accurately predicts the motion trajectory of fruit under wind disturbances, resolving the issue of inaccurate dynamic target tracking using traditional methods. By comprehensively considering multiple key factors such as fruit occlusion rate, growth posture, depth value, motion state and complexity of the surrounding environment, a scientific picking suitability scoring function was constructed; the working parameters of the end effector were automatically adjusted according to different picking difficulties, overcoming the limitation of single picking point selection in existing technology.

[0034] In a specific embodiment, before executing step S101, the following steps are further included:

[0035] The parameters of the large-viewing-area sensor are set so that the field of view angle of the large-viewing-area sensor is greater than 120°, the resolution reaches 1920×1080 pixels, and the acquisition frequency is set to 30 frames / second. The parameters of the small-viewing-area sensor are set so that the field of view angle of the small-viewing-area sensor is controlled within the range of 30°-60°, the resolution reaches 2560×1440 pixels, and the acquisition frequency is set to 60 frames / second.

[0036] The large field of view sensor and the small field of view sensor are rigidly connected and installed, and the precise positional relationship between the sensor and the mechanical structure is determined through a hand-eye calibration algorithm to obtain the installation positions of the large field of view sensor and the small field of view sensor;

[0037] Based on the installation positions of the large field of view sensor and the small field of view sensor, the robot body base coordinate system, the robot picking mechanism coordinate system, the large field of view sensor coordinate system, the large field of view sensor target picking preparation point coordinate system, the small field of view sensor coordinate system and the small field of view sensor target picking point coordinate system are established to obtain a cascaded vision system.

[0038] Specifically, a large-field-of-view sensor suitable for macroscopic perception was selected based on the spatial scope of the orchard's operating environment and the visual coverage requirements of the harvesting task. This sensor required a wide viewing angle and medium resolution to complete the initial detection and location of all fruit in the scene. Its field of view was set to greater than 120°, enabling simultaneous coverage of a large orchard area without requiring frequent viewing angle shifts. Its resolution was set to 1920×1080 pixels, ensuring image clarity while controlling data volume for real-time processing. The acquisition frequency was set to 30 frames per second, ensuring sufficient image frames to support continuous tracking of fruit during low- and medium-speed movement. The small field of view sensor used in conjunction with the large and small field of view sensor focuses on localized, high-precision imaging of individual fruit areas. Its design principles prioritize both resolution and timeliness. Its field of view is controlled within a range of 30° to 60°, narrowing the image range and enhancing target focus. The resolution is set to 2560×1440 pixels, enabling the system to accurately capture details such as fruit edges and stem connections. The acquisition frequency is increased to 60 frames per second, enabling stable capture of high-speed moving targets and supporting subsequent pixel-level mask segmentation and pose extraction. After sensor parameter configuration, the large and small field of view sensors are rigidly mounted on the same end module of the harvesting robot arm. This structural connection is torsionally and vibrationally resistant, and prevents relative displacement, ensuring a fixed spatial relationship between the two sensors during operation and preventing visual mismatches caused by vibration or external forces. After installation, a hand-eye calibration algorithm is used to determine the spatial transformation between the sensor and the robot's mechanical structure. This process utilizes Zhang's calibration method coupled with a nonlinear least-squares optimization algorithm. A calibration plate containing high-contrast landmarks is set up, and at least twenty image sequences are captured at various angles and distances. Image processing algorithms are used to extract pixel coordinates of corner points. Combined with the real-time pose data of the picking mechanism, a mapping function is established from camera image coordinates to machine coordinates. During the optimization process, the system calculates rotation and translation vectors and combines them into a rigid body transformation matrix. This matrix is ​​then adjusted using reprojection error as a calibration accuracy metric to obtain the precise three-dimensional installation positions and attitude orientation parameters of the two sensors in the robot coordinate system. To achieve coordination between multiple visual sources and a complete mapping from visual perception to motion control, the system establishes a standardized multi-coordinate system after acquiring position parameters. This coordinate system uses the robot's base coordinate system as the global reference frame. Six sub-coordinate systems are constructed based on this system: the robot picking mechanism coordinate system, the large field of view sensor coordinate system, the large field of view target picking preparation point coordinate system, the small field of view sensor coordinate system, and the small field of view target picking point coordinate system. These coordinate systems are interconnected through homogeneous transformations.The robot's base coordinate system describes the position and posture of the entire picking platform, the robot's picking mechanism coordinate system describes the geometric motion structure of the manipulator and end effector, and the large-view sensor coordinate system corresponds to its imaging position and orientation. The large-view target picking point coordinate system represents the spatial position of the fruit after macroscopic initial positioning. The small-view sensor coordinate system represents the angle and coordinate range of local image acquisition, and the small-view target picking point coordinate system represents the fine spatial position used to execute the final picking action. The transformation relationship between these coordinate systems is encapsulated using rigid body rotation matrices and translation vectors, forming a coordinate mapping chain. The initial fruit position acquired in the large-view sensor coordinate system is first transformed from the large-view sensor coordinate system to the robot's base coordinate system, and then translated by the robot control system into the picking mechanism coordinate system to guide the small-view sensor to move within the viewing range of the picking preparation point. The target picking point acquired in the small-view coordinate system is then reversely transformed from the small-view coordinate system to the robot's execution layer coordinate system to control the gripping trajectory and angular posture of the end effector gripper. In the entire cascade vision system, the large field of view provides large-scale fruit retrieval and distribution mapping functions, while the small field of view is responsible for fine identification and dynamic tracking. The two are interconnected through a coordinate system and rely on accurate rigid installation and calibration technology to achieve unified expression of spatial information and consistent temporal control, forming a highly robust multi-level vision fusion system that can operate stably in an orchard environment.

[0039] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0040] The original image data is collected by a large field of view sensor in a cascaded vision system, and the original image data is subjected to grid division processing to obtain a set of divided image sub-regions;

[0041] The divided image sub-region set is input into the adaptive control model of the encoder-decoder architecture, and feature encoding is performed through the 5-layer convolutional structure in the encoder to obtain image coding features;

[0042] The image encoding features are decoded through the 5-layer transposed convolution structure in the decoder to obtain the initial decoding result;

[0043] The initial decoding result is input into the first processing branch to perform pixel-level enhancement to obtain pixel-level enhancement features, and the initial decoding result is input into the second processing branch to perform ISP parameter learning to obtain global adjustment parameters;

[0044] The pixel-level enhancement features are fused with the global adjustment parameters to obtain the enhanced large-viewing-area image.

[0045] Specifically, by collecting raw image data through the large-field-of-view sensor in the cascaded vision system, the image covers a wide range and can capture multiple fruit targets and their natural background environment at once, including typical interference elements such as direct sunlight, leaf occlusion, and crossing fruit branches. To improve the adaptability to local areas, the collected raw image is rasterized and divided according to preset rules. The entire image is equidistantly divided into several sets of image sub-regions of uniform size and with a certain overlap. Each sub-region covers several local segments of the fruit and its environmental background. The divided overlapping areas can effectively reduce the problem of fragmentation of boundary information and provide sufficient contextual cross-features for the network model. The divided image sub-region sets are sequentially input into the image adaptive control model built based on the encoder-decoder architecture. In the encoder part of this model, feature encoding operations are performed on each image sub-region through a 5-layer convolutional structure. This five-layer convolutional architecture uses a standard convolution kernel size, typically 3×3, and applies batch normalization and nonlinear activation functions after each layer to maintain stability and nonlinear representation during feature extraction. The number of channels increases from shallow to deep layers, enabling the network to gradually compress the spatial dimensions of the image while enhancing its ability to perceive semantic features. This allows the network to extract intermediate feature maps that contain information such as illumination variations, edge structure, and color distribution. The encoder outputs the image's encoded features. The decoder then decodes these encoded features using a five-layer transposed convolutional architecture. Upsampling gradually restores the image's spatial dimensions while reconstructing an intermediate image representation that retains the original image's structure. Each layer of transposed convolution is combined with a skip connection mechanism to fuse shallow spatial details with deep semantic information at the same feature scale, maximally preserving fine-grained details such as the original image's texture, edges, and contours. After the five-layer decoding is complete, an initial decoded image with intact spatial structure is obtained. This initial decoded image then enters a two-way parallel processing architecture for pixel-level enhancement and global parameter learning. The first processing branch focuses on pixel-level image enhancement. It incorporates three serially connected pixel enhancement modules, each of which incorporates a residual structure, dilated convolution, and an attention mechanism to enhance image recovery in the presence of blurred boundaries, dark areas, and texture loss. The residual structure helps the model learn incremental image changes, while dilated convolution expands the receptive field to capture long-range information. The attention mechanism adaptively adjusts the enhancement intensity of each region, giving higher response weights to areas of actual fruit in the image. After processing, this branch outputs a pixel-level enhanced image feature map. Simultaneously, the second processing branch performs ISP parameter learning for the entire image, adaptively estimating image signal processing parameters. These parameters primarily encompass brightness balance, color correction, and gamma adjustment.A lightweight fully connected network structure is constructed within this branch. The input is the global feature map of the initial decoded image. A set of matrix coefficients and brightness adjustment function curve parameters that control the image color conversion are calculated through a set of attention maps to form a set of global image adjustment parameters. The goal of this processing path is to unify the color style and contrast range of the entire image, ensure that images in different areas and under different lighting conditions have consistent visual characteristics, and provide a more stable input distribution for subsequent detectors. After the two branches complete the processing respectively, the pixel-level enhanced features are fused with the global adjustment parameters, including weighted integration in the image space and mapping inverse calculation in the feature space, and finally output an enhanced large-viewing-area image.

[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0047] The enhanced large-viewing-area image is input into the target detection network for depthwise separable convolution processing to obtain the initial feature map;

[0048] The CBAM attention module performs channel attention and spatial attention calculations on the initial feature map to obtain an enhanced feature map.

[0049] Through three detection heads of different scales, multi-scale feature fusion detection is performed on the enhanced feature map to obtain the location information of the fruit target;

[0050] Perform non-maximum suppression on the position information to obtain an accurate bounding box of the fruit target;

[0051] The center point coordinates of each fruit target are calculated based on the precise fruit target bounding box as the coordinates of the picking preparation point. At the same time, the depth value of the corresponding point is obtained from the large-viewing area sensor to obtain the position of the picking preparation point.

[0052] Execute the coordinate transformation matrix on the position of the picking preparation point, and transform the picking preparation point position from the large field of view sensor coordinate system to the robot body base coordinate system, and then transform it to the robot picking mechanism coordinate system to obtain the position parameters that control the movement of the small field of view sensor;

[0053] Based on the position parameters for controlling the movement of the small viewing area sensor, the small viewing area sensor is guided to the picking preparation point position.

[0054] Specifically, the enhanced large-viewing-area image is fed into an object detection network for feature extraction and spatial localization. This object detection network uses a depthwise separable convolutional architecture as its backbone. Depthwise separable convolution decomposes standard convolution into two stages: channel-wise and point-wise convolution. These stages capture intra-channel features and cross-channel interactions, respectively, improving the model's ability to recognize fruit outlines, color boundaries, and background structures in the image. This module generates an initial feature map that preserves the information distribution and contour patterns of key image regions. Next, the CBAM attention module performs channel-wise and spatial-attention calculations on the initial feature map. In the channel-wise attention stage, average pooling and max pooling are performed on the initial feature map in the spatial dimension to extract a global description vector. Channel-wise weighting coefficients are generated using a multi-layer perceptron (MLP) to enhance the response of the fruit channel and suppress irrelevant background features. In the spatial-attention stage, channel-wise pooling and convolution are performed to generate a spatial weight map, allowing the network to focus on highly responsive regions of the image, such as fruit edges, stalk connections, and color transitions. The CBAM module outputs the enhanced feature map. The enhanced feature maps are then subjected to multi-scale feature fusion detection using three detection heads of different scales. These heads correspond to the receptive regions of small, medium, and large objects, respectively, comprehensively covering the various sizes and scales of fruit that appear in the image. Through operations such as convolution, normalization, and activation, the detection heads output the fruit's category probability, bounding box location, and confidence information. Each detection head processes a set of downsampled feature maps and extracts candidate objects within a specific scale range. This enables the system to maintain high recognition consistency and coverage even in orchard images with varying fruit sizes and uneven distribution from near to far. To remove duplicate candidate bounding boxes generated during multi-scale detection, a non-maximum suppression mechanism is introduced based on the object detection results. This mechanism ranks candidate boxes of the same category with high overlap by confidence and sequentially removes low-confidence bounding boxes whose overlap exceeds a set threshold. This retains the unique, optimal bounding box for each fruit, uniquely identifying boundaries and accurately extracting edges. The resulting dataset of fruit object bounding boxes is well-structured and accurately covers the fruit. The coordinates of the center point of the fruit target are calculated based on each precise bounding box. This center point is defined as the fruit's pre-harvest point and serves as a geometric reference for subsequent fine alignment. To map this center point from the image coordinate system to real-world coordinates, the depth value corresponding to this center point is extracted from the large-viewing-area image, and the physical position of this point in three-dimensional space is restored based on the imaging geometry. These three-dimensional coordinates represent the spatial position of the fruit's pre-harvest point at the current moment and perspective, including the relative coordinate relationships in the horizontal, vertical, and front-to-back directions.The three-dimensional picking preparation point position undergoes multiple coordinate transformations to complete the transition from perception space to execution space. The picking point is mapped from image coordinates to the robot's reference coordinate system using a pre-calibrated transformation matrix from the large-viewing-area sensor to the robot's base coordinate system. This transformation matrix, defined in the picking mechanism's structural parameters, then transfers this information from the reference coordinate system to the picking mechanism's own working coordinate system. This generates the pose control parameters for the joint movement of the picking end point and the small-viewing-area sensor. These control parameters include the target position and motion planning information, such as the desired arrival posture and approach angle, ensuring a smooth subsequent movement path and a reasonable gripping posture. These pose control parameters are then transmitted to the picking mechanism's motion control unit, guiding the small-viewing-area sensor to precisely move along the planned trajectory to the target picking preparation point. During this process, closed-loop corrections are performed in conjunction with the end point's position feedback signal to improve alignment accuracy and adjust the small-viewing-area sensor's focus and fixed viewing angle, ensuring that its final field of view fully encompasses the target fruit area and its surrounding structures.

[0055] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0056] The local image data is collected by the small field of view sensor in the cascade vision system, and the local image data is input into the multi-object segmentation network for pixel-level mask segmentation to obtain the initial pixel-level mask;

[0057] The ResNext-101-FPN structure is used to extract multi-scale feature maps from the initial pixel-level mask to obtain multi-scale feature representation;

[0058] Perform bounding box regression and classification calculations on the multi-scale feature representation to obtain the precise bounding box of the fruit target, and generate pixel-level masks based on the precise bounding box of the fruit target to obtain the precise segmentation mask of the fruit target;

[0059] The occlusion rate of the fruit target is calculated based on the precise segmentation mask, and the angle between the fruit posture angle and the vertical direction is calculated based on the precise segmentation mask to obtain the growth posture information of the fruit target.

[0060] Specifically, local image data is collected using a small-field-of-view sensor within a cascaded vision system. Due to its narrow field of view, the small-field-of-view sensor can focus on the target fruit region at a more concentrated imaging angle, capturing local image data that includes details of the fruit structure, the position of the stalk, and information about occlusions from surrounding leaves. This local image data is then fed into a multi-object instance segmentation network for pixel-level mask segmentation. This network, designed to integrate the dual capabilities of instance detection and semantic segmentation, can both identify the categories of different target objects in an image and distinguish instances of objects within the same category. Within each target region, a pixel-level binary mask with the same scale as the original image is generated, resulting in an initial mask result that covers the main outline of the fruit, the edge transition zone, and some occluded but predictable areas. To improve the adaptability and accuracy of the mask segmentation results for fruits of varying scales and structural characteristics, the ResNext-101 architecture is used as the backbone based on the initial mask result, combined with a feature pyramid network for multi-scale feature map extraction. ResNext-101 achieves parallel feature extraction within the channel through grouped convolution, significantly reducing the computational complexity while maintaining high network expression capabilities, and has good representation capabilities for complex structures such as texture, color blocks, edges, and fruit stem connections on the fruit surface. After combining the FPN structure, feature maps at multiple scale levels are extracted from the image, such as the original Figure 1The system generates feature maps with scaling ratios of 1 / 4, 1 / 8, and 1 / 16. Low-level feature maps retain more edge details and texture information, while high-level feature maps extract the overall semantic outline and spatial consistency of the fruit. Through horizontal fusion and upsampling, a fused feature representation containing multi-scale semantic features is formed, resulting in a structurally stable and semantically rich multi-scale feature map. Based on this multi-scale feature map representation, bounding box regression and object classification operations in the object detection process are performed. A convolutional module predicts the bounding box position, width and height, and object category of each candidate region. Redundant candidate boxes are processed through non-maximum suppression to output precise bounding box data that closely matches the fruit location. Each bounding box defines a region of interest. The system then activates an instance mask generation module within the bounding box to reconstruct a pixel-level mask based on the local feature map. This process utilizes a multi-layer transposed convolutional architecture with skip connections. Through progressive upsampling and fine decoding, object instance masks are generated that match the original image resolution. The mask boundaries not only maintain clear spatial continuity but also flexibly fit the transition between the fruit edge and the background, resulting in a structurally complete and precisely defined fruit segmentation mask. Based on the precise segmentation mask calculation, geometric and area analysis is performed. By counting the total number of pixels marked as fruit in the mask image and comparing the actually visible pixel area with the occluded or incomplete areas, the visible proportion of the fruit within the current image viewing angle is assessed. This ratio can be used to infer the degree of occlusion of the fruit, forming an occlusion ratio indicator. A higher occlusion ratio indicates that the fruit is more obscured by surrounding leaves, branches, or other fruit, making it more difficult to identify and grasp. A lower occlusion ratio indicates that the object has clear boundaries and intact structure, making it a preferred target for picking. Contour extraction is also performed based on the edge coordinates of the precise mask image to construct a set of boundary points for the fruit's outline, from which the principal morphological directions are calculated. Principal component analysis is performed on the boundary point set, extracting the first principal axis direction from the set as the fruit's growth orientation. The angle between this principal direction vector and the image vertical is then calculated as the fruit's growth attitude angle. The attitude angle reflects the fruit's tilt or deviation from its growth direction. This angle is used to guide the end effector's approach path adjustment and the gripper's attitude control in spatial control. The occlusion rate and growth attitude angle are structured perception indicators used to evaluate whether the fruit has good recognition clarity and picking operability. A high occlusion rate or a large attitude angle will lead to operational risks such as recognition failure or clamping and falling off. Therefore, these two indicators are used as key inputs in the subsequent picking suitability scoring model to evaluate the picking priority of the target fruit.

[0061] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0062] Based on the depth information, RGB-D information is obtained, a depth value is assigned to each pixel, and the two-dimensional image coordinates are converted into three-dimensional space coordinates by combining the camera intrinsic parameter matrix to obtain pseudo point cloud data of the fruit area;

[0063] Based on the pseudo point cloud data of the fruit area, a horizontal XOY coordinate system is established with the picking mechanism as the origin, and the three-dimensional position of the fruit is projected onto the horizontal plane to form a BEV coordinate representation;

[0064] The segmentation results of the small field of view sensor are used to extract the coordinates of the center point of the fruit bounding box and the corresponding depth value, and the three-dimensional spatial position of the fruit target is calculated by combining the BEV coordinate representation;

[0065] The three-dimensional spatial position of the fruit target is matched between frames to obtain the motion trajectory of the fruit target, and the Kalman filter algorithm is used to predict the state of the motion trajectory of the fruit target to obtain the fruit state prediction data;

[0066] An adaptive process noise adjustment mechanism is introduced based on the fruit state prediction data, and the process noise covariance matrix is ​​dynamically adjusted according to the change of fruit position in adjacent frames to obtain motion trajectory prediction data.

[0067] Specifically, relying on the depth perception capabilities of the small-field-of-view sensor in the cascaded vision system, an RGB-D image containing color image information and spatial depth data is acquired. Each pixel has a color value composed of RGB channels and a corresponding depth value, which represents the distance between the pixel and the camera at the time of imaging. To transform the fruit information in the two-dimensional image into an operational geometric entity in three-dimensional space, the depth value is combined with the camera's intrinsic parameter matrix. The focal length and principal point coordinates in the intrinsic parameters are used to perform back-projection processing on each pixel in the image, thereby achieving a mapping of the two-dimensional image coordinates to the three-dimensional world coordinates. In this process, each pixel is assigned a spatial position, which is aggregated in the fruit recognition area to form a dense set of points with spatial depth attributes, namely the pseudo point cloud data of the fruit area. Based on pseudo-point cloud data from the fruit region, a horizontal XOY coordinate system parallel to the ground is established, with the location of the picking mechanism's end-effector as the coordinate origin. Within this coordinate system, the fruit's three-dimensional position is converted to a BEV representation through projection. This representation method projects the three-dimensional points in the pseudo-point cloud onto the XOY plane, thereby obtaining the fruit's planar position distribution from a bird's-eye view. This allows the system to intuitively represent the fruit's dynamic trajectory on the ground coordinates and facilitates operations such as trajectory fitting, motion trend determination, and path optimization. Based on this coordinate system, the relative displacement of the fruit target on the BEV plane is continuously monitored over subsequent time series, thereby constructing a dataset of the fruit's planar path over time. Simultaneously, after being guided to the target fruit picking preparatory point, the small-field-of-view sensor performs high-resolution imaging of the fruit and pixel-level instance segmentation. The system extracts the fruit's bounding box center from the small-field-of-view segmentation result as the target reference coordinate in the local image. Simultaneously, the depth value corresponding to this center point is extracted from the sensor's depth channel to obtain a 3D representation of its position in the local coordinate system. These three-dimensional coordinates are geometrically aligned with the BEV coordinate system, completing the spatial fusion of information between the two types of visual sensors. This enables the system to track the fruit spatially with multi-resolution and multi-angle views, thereby improving tracking accuracy and matching stability. The three-dimensional spatial position of the fruit target is matched between frames. By calculating the Euclidean distance, relative direction, and velocity change of the fruit position at consecutive moments, it is determined whether the target in the current frame is a continuation of the target in the previous frame. A target ID tracking strategy is used to maintain the temporal consistency of each fruit. After successful matching, a chronological trajectory of the fruit in three-dimensional space is constructed, reflecting the dynamic swaying behavior of the fruit in the natural environment caused by wind disturbance, gravity sway, or branch elasticity. To improve the robustness of the fruit trajectory analysis, a Kalman filter algorithm is introduced to predict the state of the above-mentioned trajectory. This filter jointly models the current three-dimensional position and velocity state of the fruit. At each frame, the fruit position is predicted based on the state of the previous frame and fused with the current observed position to produce a smooth and continuous state estimate.The Kalman filter process consists of two parts: a state transfer equation and an observation update equation. The state transfer equation assumes the fruit moves over a small range at a constant velocity or slowly varying acceleration, while the observation update equation incorporates the actual measured position in the current frame. Bayesian estimation theory is used to optimally fuse the states, resulting in a smooth path representation that more closely approximates the true trajectory. This is suitable for maintaining stable recognition and tracking of fruit in frames partially obscured by leaves or blurred by distortion. Considering that fruit motion in natural scenes does not strictly adhere to the linear system assumption, an adaptive process noise adjustment mechanism is incorporated into the Kalman filter framework. This mechanism dynamically adjusts the process noise covariance matrix based on the change in fruit position between adjacent frames. When the system detects dramatic changes in fruit position or significant oscillation, the process noise weight is automatically increased to expand the prediction uncertainty range to accommodate the sudden changes in real-world motion. Conversely, when the fruit position in consecutive frames is relatively stable or exhibits a periodic pattern, the noise covariance is reduced to enhance prediction accuracy and stability. The adaptive adjustment mechanism driven by displacement changes makes the filter environmentally adaptive, dynamically adjusting its prediction model under different fruit tree shapes, wind speed conditions, and picking path disturbances, thereby maintaining efficient capture of the fruit's true motion state and obtaining motion trajectory prediction data.

[0068] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0069] A fruit picking factor set is created based on the occlusion rate, the posture angle in the growth posture information, the depth value in the motion trajectory prediction data, the motion state and the complexity of the surrounding environment;

[0070] A picking suitability scoring function is constructed based on the fruit picking factor set, and the fruit picking factor set is scored using a scoring calculation model to obtain a picking suitability score.

[0071] Make picking decisions based on the picking suitability scores to obtain a set of fruits to be picked;

[0072] For each fruit in the set of fruits to be picked, the position of the fruit stem connection point is determined based on the segmentation results of the small-viewing-area sensor, and the optimal picking point is calculated.

[0073] Generate picking execution instructions based on the optimal picking point and picking suitability score.

[0074] Specifically, a set of fruit picking factors is created based on the occlusion rate, attitude angle from growth posture information, depth value from motion trajectory prediction data, motion state, and surrounding environment complexity. The occlusion rate is derived from the fruit mask generated by the small-view segmentation network. The target clarity is quantified by analyzing the area ratio of the fruit area to the occluded area. The attitude angle is calculated by the angle between the main contour direction and the image perpendicular, reflecting the fruit's tilt and the complexity of its growth posture. The depth value is derived by fusing the large-view RGB-D image with the depth information of the center pixel of the small-view center and is an important geometric parameter for assessing target distance. The motion state is analyzed based on the velocity and acceleration state vectors output by the Kalman filter prediction model to reflect the stability and trackability of the fruit under external forces. The environmental complexity comprehensively considers the leaf density, branch occlusion, lighting contrast, and background interference characteristics around the target. The environmental vision module extracts feature maps of adjacent regions and performs cluster analysis and complexity scoring modeling to represent the external interference level of the current target. A picking suitability scoring function is constructed based on the fruit picking factor set. This scoring function is derived through parameter settings based on the experience of picking experts and regression modeling of historical picking experimental data. Each factor in the function is assigned a different weight coefficient to reflect its impact on the actual picking success rate. For example, occlusion rate and posture angle are given higher weights because they directly affect the feasibility of visual recognition and mechanical gripping. Depth value and motion state are medium-weighted parameters, used to limit the accessibility and stability of the picking path. Environmental complexity is an auxiliary weight used for risk estimation and operation time assessment. In the actual scoring process, the factor set for each fruit is input into the scoring function to calculate a unified picking suitability score. This score numerically reflects whether the current target is ready for picking and forms a sortable priority system among multiple targets. Once the picking suitability scores for all candidate target fruits are calculated, a threshold strategy is set to execute the picking decision. Fruits with scores exceeding a certain threshold are marked as targets for picking, and these targets are unified to form a set of fruits for picking. This set serves as the input for subsequent path planning, end-gripping trajectory generation, and picking action scheduling. Targets are ranked from highest to lowest according to their scores, ensuring priority execution for fruits with high success rates, low losses, and low path costs. For each target in the set of fruits to be picked, the high-resolution mask output by the small-view segmentation network is reused. By analyzing the fruit's contour points and edge structure, the connecting region between the fruit and the stem is extracted. This region appears as a slender transition zone between the upper contour of the fruit and the background, with distinct characteristics in terms of edge gradient, color change, and geometric shape.The position of the fruit stalk connection point is determined through edge intensity gradient analysis and shape prior discrimination algorithm, and a local coordinate system of the fruit is established near this point. In this coordinate system, the direction of the fruit stalk is used as the z-axis direction, and the position of the fruit center of mass and the local surface normal vector are combined to calculate the optimal picking point, that is, the spatial contact position that can minimize clamping interference, ensure the picking angle and avoid obstructions. This position not only takes into account the shortest approach path of the robot arm, but also integrates the geometric symmetry of the fruit and the structural stability of the fruit stalk to ensure that the fruit does not fall off, break or twist during the clamping process. The optimal picking point is integrated with the picking suitability score corresponding to the target fruit to generate a standardized picking execution instruction. The instruction contains the position coordinates and posture direction parameters that the robot arm needs to reach, as well as control variables such as grasping speed, clamping force, and action time window. The scoring result is used to decide whether to adopt a low-speed, high-precision clamping mode (such as if the score is on the critical edge) or a high-speed, fast operation mode (such as if the score is extremely high). For fruits with a score below the threshold but slightly above the abandonment boundary, the system determines whether their suitability can be improved after a certain posture adjustment. If so, they enter the waiting window queue. Otherwise, picking is postponed and re-evaluated in the subsequent detection cycle.

[0075] The above describes the target fruit dynamic tracking method based on visual perception and deep learning in the embodiment of the present invention. The following describes the target fruit dynamic tracking system based on visual perception and deep learning in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a target fruit dynamic tracking system based on visual perception and deep learning includes:

[0076] The image enhancement module 201 is used to collect original image data and perform image enhancement through an adaptive control model of an encoder-decoder architecture to obtain an enhanced large-viewing-area image;

[0077] The target detection module 202 is used to perform deep learning target detection on the enhanced large-viewing-area image, extract the location information of the fruit target and the location of the picking preparation point, and guide the preset small-viewing-area sensor to the location of the picking preparation point;

[0078] Pixel-level segmentation module 203, used to collect local image data and perform pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit object;

[0079] A dynamic tracking module 204 is configured to dynamically track the fruit target based on the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data;

[0080] The calculation module 205 is used to calculate the picking suitability score and the optimal picking point according to the occlusion rate, the growth posture information and the motion trajectory prediction data, and generate a picking execution instruction.

[0081] Through the collaborative operation of these components, the wide-field-of-view (FOV) and narrow-field-of-view (FOV) sensors work together to overcome the technical bottleneck of a single visual sensor's inability to simultaneously achieve global perception and localized accurate recognition. This system achieves multi-scale, multi-angle, and multi-resolution collaborative perception, significantly improving fruit detection coverage and positioning accuracy in complex orchard environments. An adaptive image manipulation model employing an encoder-decoder architecture effectively mitigates the impact of image interference factors such as strong light, shadows, occlusions, and motion blur on image quality in orchard environments. Two dedicated processing branches, respectively for pixel-level enhancement and global ISP parameter learning, enable the system to adapt to varying lighting conditions. A deep learning-based object detection and multi-object segmentation network enhances fruit feature extraction. A multi-object instance segmentation network based on the MASNet architecture effectively identifies partially occluded fruit, providing accurate occlusion rate and fruit pose information for subsequent picking decisions. Combining RGB-D information to construct pseudo point cloud data and BEV views enables fruit localization in three-dimensional space. A Kalman filter algorithm, incorporating an adaptive process noise adjustment mechanism, accurately predicts fruit motion under wind disturbances, addressing the inaccuracy of traditional dynamic object tracking methods. By comprehensively considering multiple key factors such as fruit occlusion rate, growth posture, depth value, motion state and complexity of the surrounding environment, a scientific picking suitability scoring function was constructed; the working parameters of the end effector were automatically adjusted according to different picking difficulties, overcoming the limitation of single picking point selection in existing technology.

[0082] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for enabling a target fruit dynamic tracking device based on visual perception and deep learning (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program code.

[0084] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target fruit dynamic tracking method based on visual perception and deep learning, characterized in that: include: Collect the original image data and perform image enhancement through the adaptive control model of the encoder-decoder architecture to obtain the enhanced large-viewing-area image; Performing deep learning target detection on the enhanced large-viewing-area image, extracting the location information of the fruit target and the location of the picking preparation point, and guiding the preset small-viewing-area sensor to the location of the picking preparation point; Collect local image data and perform pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit target; Dynamically tracking the fruit target according to the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data; According to the occlusion rate, the growth posture information and the motion trajectory prediction data, a picking suitability score and an optimal picking point are calculated, and a picking execution instruction is generated.

2. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 1 is characterized in that, Before acquiring the original image data, it also includes: The parameters of the large-viewing-area sensor are set so that the field of view angle of the large-viewing-area sensor is greater than 120°, the resolution reaches 1920×1080 pixels, and the acquisition frequency is set to 30 frames / second. The parameters of the small-viewing-area sensor are set so that the field of view angle of the small-viewing-area sensor is controlled within the range of 30°-60°, the resolution reaches 2560×1440 pixels, and the acquisition frequency is set to 60 frames / second. The large viewing area sensor and the small viewing area sensor are rigidly connected and installed, and a precise positional relationship between the sensors and the mechanical structure is determined using a hand-eye calibration algorithm to obtain the installation positions of the large viewing area sensor and the small viewing area sensor; Based on the installation positions of the large field of view sensor and the small field of view sensor, the robot body base coordinate system, the robot picking mechanism coordinate system, the large field of view sensor coordinate system, the large field of view sensor target picking preparation point coordinate system, the small field of view sensor coordinate system and the small field of view sensor target picking point coordinate system are established to obtain a cascaded vision system.

3. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 2 is characterized in that, The method collects original image data and performs image enhancement through an adaptive control model of an encoder-decoder architecture to obtain an enhanced large-viewing-area image, including: Collecting original image data through the large viewing area sensor in the cascade vision system, and performing grid division processing on the original image data to obtain a divided image sub-region set; Inputting the divided image sub-region set into the adaptive control model of the encoder-decoder architecture, and performing feature encoding through the 5-layer convolutional structure in the encoder to obtain image coding features; Decoding the image encoding features through a 5-layer transposed convolution structure in the decoder to obtain an initial decoding result; Inputting the initial decoding result into the first processing branch to perform pixel-level enhancement to obtain pixel-level enhancement features, and inputting the initial decoding result into the second processing branch to perform ISP parameter learning to obtain global adjustment parameters; The pixel-level enhancement feature is fused with the global adjustment parameter to obtain an enhanced large-viewing-area image.

4. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 2 is characterized in that, The method of performing deep learning target detection on the enhanced large-viewing-area image, extracting the position information of the fruit target and the position of the picking preparation point, and guiding the small-viewing-area sensor to the position of the picking preparation point includes: Inputting the enhanced large-viewing-area image into the target detection network for depthwise separable convolution processing to obtain an initial feature map; Perform channel attention and spatial attention calculations on the initial feature map through the CBAM attention module to obtain an enhanced feature map; Through three detection heads of different scales, multi-scale feature fusion detection is performed on the enhanced feature map to obtain the position information of the fruit target; Performing non-maximum suppression processing on the position information to obtain an accurate fruit target bounding box; Calculating the center point coordinates of each fruit target based on the precise fruit target bounding box as the coordinates of the picking preparation point, and simultaneously acquiring the depth value of the corresponding point from the large viewing area sensor to obtain the position of the picking preparation point; Execute a coordinate transformation matrix on the position of the picking preparation point, sequentially transform the position of the picking preparation point from the large field of view sensor coordinate system to the robot body base coordinate system, and then transform it to the robot picking mechanism coordinate system, to obtain position parameters for controlling the movement of the small field of view sensor; Based on the position parameters for controlling the movement of the small viewing area sensor, the small viewing area sensor is guided to the picking preparation point position.

5. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 4 is characterized in that, The method collects local image data and performs pixel-level segmentation through a multi-target segmentation network to obtain the occlusion rate and growth posture information of the fruit target, including: Collecting local image data through a small field of view sensor in the cascaded vision system, and inputting the local image data into a multi-object segmentation network for pixel-level mask segmentation to obtain an initial pixel-level mask; Using the ResNext-101-FPN structure, a multi-scale feature map is extracted from the initial pixel-level mask to obtain a multi-scale feature representation; Performing bounding box regression and classification calculation on the multi-scale feature representation to obtain an accurate bounding box of the fruit target, and performing pixel-level mask generation on the multi-scale feature representation based on the accurate bounding box of the fruit target to obtain an accurate segmentation mask of the fruit target; The occlusion rate of the fruit target is calculated based on the precise segmentation mask, and the angle between the fruit posture angle and the vertical direction is calculated based on the precise segmentation mask to obtain the growth posture information of the fruit target.

6. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 1 is characterized in that, The method of dynamically tracking the fruit target based on the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data includes: Acquire RGB-D information based on the depth information, assign a depth value to each pixel, and convert the two-dimensional image coordinates into three-dimensional space coordinates in combination with the camera intrinsic parameter matrix to obtain pseudo point cloud data of the fruit area; Based on the pseudo point cloud data of the fruit area, a horizontal XOY coordinate system is established with the picking mechanism as the origin, and the three-dimensional position of the fruit is projected onto the horizontal plane to form a BEV coordinate representation; Extracting the coordinates of the center point of the fruit bounding box and the corresponding depth value from the segmentation result of the small viewing area sensor, and calculating the three-dimensional spatial position of the fruit target in combination with the BEV coordinate representation; Performing inter-frame matching on the three-dimensional spatial position of the fruit target to obtain a motion trajectory of the fruit target, and using a Kalman filter algorithm to perform state prediction on the motion trajectory of the fruit target to obtain fruit state prediction data; An adaptive process noise adjustment mechanism is introduced based on the fruit state prediction data, and the process noise covariance matrix is ​​dynamically adjusted according to the change in the fruit position of adjacent frames to obtain motion trajectory prediction data.

7. The target fruit dynamic tracking method based on visual perception and deep learning according to claim 1 is characterized in that, The step of calculating a picking suitability score and an optimal picking point based on the occlusion rate, the growth posture information, and the motion trajectory prediction data, and generating a picking execution instruction, includes: Creating a fruit picking factor set based on the occlusion rate, the posture angle in the growth posture information, the depth value in the motion trajectory prediction data, the motion state and the complexity of the surrounding environment; Constructing a picking suitability scoring function based on the fruit picking factor set, and scoring the fruit picking factor set using a scoring calculation model to obtain a picking suitability score; Making a picking decision based on the picking suitability score to obtain a set of fruits to be picked; For each fruit in the set of fruits to be picked, determining the position of the fruit stem connection point according to the segmentation result of the small viewing area sensor, and calculating the optimal picking point; A picking execution instruction is generated according to the optimal picking point and the picking suitability score.

8. A target fruit dynamic tracking system based on visual perception and deep learning, characterized in that: For implementing the target fruit dynamic tracking method based on visual perception and deep learning according to any one of claims 1 to 7, the target fruit dynamic tracking system based on visual perception and deep learning comprises: The image enhancement module is used to collect raw image data and perform image enhancement through an adaptive control model of the encoder-decoder architecture to obtain an enhanced large-viewing-area image; A target detection module is used to perform deep learning target detection on the enhanced large-viewing-area image, extract the location information of the fruit target and the location of the picking preparation point, and guide the preset small-viewing-area sensor to the location of the picking preparation point; The pixel-level segmentation module is used to collect local image data and perform pixel-level segmentation through a multi-object segmentation network to obtain the occlusion rate and growth posture information of the fruit target; A dynamic tracking module is used to dynamically track the fruit target based on the acquired depth information and the segmentation result of the small viewing area sensor to obtain motion trajectory prediction data; A calculation module is used to calculate the picking suitability score and the optimal picking point according to the occlusion rate, the growth posture information and the motion trajectory prediction data, and generate a picking execution instruction.

Citation Information

Patent Citations

  • Fruit recognition tracking method and system based on deep learning algorithm

    CN108564601A

  • String type fruit distributed visual active sensing method and application thereof

    CN111602517A

  • Target detection method based on multi-modal data fusion and in-vivo fruit picking method based on target detection model

    CN115376125A

  • RAW domain low-light image enhancement method and system imitating traditional ISP assembly line

    CN116739916A

  • Robot operation track dynamic planning method and device based on mixed vision

    CN116852349A

Cited By

  • Citrus picking robot control method based on double-vision cooperation and visual servo

    CN122139565A

  • Dynamic fruit detection tracking and real-time positioning method of citrus picking robot

    CN122244107A