Automatic vegetable and fruit picking and grading method based on AI vision and dynamic control
By constructing a multi-angle spectral texture feature library and viewpoint sensitivity modeling, combined with camera viewpoint arrangement and fruit identity binding mechanism, stable identification of fruit maturity and efficient harvesting are achieved. This solves the problems of misjudgment and low efficiency caused by viewpoint occlusion in existing technologies, and improves the robustness and accuracy of the system.
Patent Information
- Application Number
- CN202511269551.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, fruit maturity recognition relies on visual detection methods based on single-frame images, which leads to inconsistent maturity labels from different perspectives, causing misjudgments, fluctuations, and low harvesting efficiency.
A multi-angle spectral texture feature library is constructed to generate a viewpoint sensitivity index and visibility prediction map. Through camera viewpoint arrangement and fruit identity binding mechanism, multi-view consistency verification and maturity confidence interval generation are achieved. A closed-loop optimization process is constructed by combining a re-observation strategy.
This solves the problems of misjudgment and low efficiency caused by occlusion in fruit maturity identification, and improves the system's environmental adaptability, identification robustness and operation accuracy.
Smart Images

Figure CN121147751A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of agricultural intelligence technology, in particular to a vegetable and fruit automatic picking and grading method based on AI vision and dynamic control. BACKGROUND
[0002] The "vegetable and fruit automatic picking and grading based on AI vision and dynamic control" refers to the fusion of artificial intelligence image recognition technology and robot dynamic control mechanism to realize the accurate identification of fruit maturity, appearance quality and spatial position, and rely on intelligent algorithm to complete the whole process of automatic operation of picking path planning, fruit clamping adjustment and quality grade determination. The method loads high-precision cameras and deep learning models to obtain real-time image feature information of the fruit, identify the type, maturity state, defect condition and shielding environment of the fruit, and then drive the multi-degree-of-freedom robot arm to avoid obstacles and perform flexible picking. At the same time, the picking objects are divided into different grades according to the set standard, and can be worked in turn according to the priority of the picking strategy. Not only can it greatly replace manual picking and sorting labor, but also can be expanded to fruit auxiliary pollination, diseased fruit identification and intelligent fruit thinning, etc. to build an integrated intelligent picking system with environmental adaptability, autonomous learning ability and work expansion ability.
[0003] The prior art has the following disadvantages: In the prior art, fruit maturity recognition generally relies on visual detection methods based on single-frame images, which analyze fruit skin color, texture features and other characteristic factors to generate maturity labels. However, in actual orchard operations, the mature area of the fruit (such as the color deepening area, the sugarification patch) usually shows asymmetric distribution, and the maturity characteristics can only be revealed at a specific angle. When the picking robot switches the visual angle due to path planning, autonomous obstacle avoidance or work rhythm adjustment, the originally clear and visible maturity patch may be blocked by leaves, fruit stems or the fruit itself, causing the recognition system to fail to detect the feature area again, and thus causing the maturity label to disappear or change at different angles. This prior art defect will directly lead to inconsistent judgments of the same fruit by the system, triggering maturity state misjudgment, picking behavior shock or mechanical arm repeated action, seriously affecting picking efficiency, decision stability and system response reliability.
[0004] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The purpose of the present application is to provide a vegetable and fruit automatic picking and grading method based on AI vision and dynamic control to solve the problems in the background.
[0006] In order to achieve the above object, the present application provides the following technical scheme: a vegetable and fruit automatic picking and grading method based on AI vision and dynamic control, comprising the following steps: A multi-angle spectral texture feature library is constructed, and a view angle sensitivity index is calculated for the surface image features corresponding to different observation angles of the fruit in the multi-angle spectral texture feature library, and a visibility prediction map is generated based on the view angle sensitivity index, which is used to mark the mature areas in the fruit surface that are easily obscured under different viewing angles; Camera view angle arrangement is implemented based on the visibility prediction map, and by formulating a camera pose control sequence, the mature areas with high view angle sensitivity indexes are kept visible in a minimum obscuring state during the collection process; After the camera view angle arrangement is completed, a spatiotemporal binding mechanism for the fruit is established, a unique identity is assigned to each fruit to be identified, and the fruit image features collected under multiple viewing angles are aggregated under the unique identity, thereby generating a maturity preliminary determination label; Multi-view consistency checking is performed based on the maturity preliminary determination label, and a maturity confidence interval is generated by fusing the image features under multiple viewing angles, which is used to eliminate the recognition errors caused by single-view collection; When the stability of the maturity confidence interval does not reach a preset threshold, a re-observation strategy is executed, the control decisions and action timing of picking are updated by fine-tuning the camera pose and combining the current mechanical arm's gripping operation window calculation results; After the picking action is completed, the picking result is fed back to the multi-angle spectral texture feature library, the view angle sensitivity index is corrected according to the residual error between the picking result and the prediction result, a feature library dynamic updating mechanism based on picking effect driving is formed, and a closed-loop optimization process of recognition and control is constructed.
[0007] Preferably, the visibility prediction map generation step is as follows: A multi-angle image collection operation is performed, and color cameras and near-infrared cameras installed on an electric turntable with rotation and pitch functions are used to collect image data at not less than twelve angles around the fixed position of the fruit sphere under constant light intensity environmental conditions, wherein each image is labeled with corresponding azimuth and pitch angle information; The color distribution information, color spot boundary features and reflection intensity distribution in the near-infrared band extracted from each image are respectively corresponded to specific physical regions in the fruit three-dimensional surface model, and on this basis, the frequency of each fruit surface region being effectively observed under each viewing angle and the number of times the corresponding key maturity features appear in the observation process are counted; According to the observation frequency and feature variation degree, the view angle sensitivity index of each surface region is calculated, and a two-dimensional visibility prediction map is generated in a polar coordinate expansion manner, and the regions with a view angle sensitivity index higher than 0.7 are extracted and labeled with a closed boundary space. Based on the visibility prediction map, the visibility safety boundary of the fruit surface is constructed, and a multi-angle collection redundancy strategy is formulated to ensure that each mature area is stably imaged at least in three low-sensitive viewing angles, meeting the determination redundancy requirement of subsequent identification.
[0008] Preferably, the steps of camera viewing angle arrangement based on the visibility prediction map are as follows: According to the generated visibility prediction map, image regions with a viewing angle sensitivity index higher than 0.7 are extracted, and the corresponding pixel positions are mapped to the three-dimensional spherical model of the fruit to obtain the spatial coordinates and surface normal vectors of each high-sensitivity region; The surface normal vector of each high-sensitivity region is used to construct a conical observation space in reverse, a plurality of preset shooting postures are set, and effective observation postures with an angle less than 45 degrees with the surface normal vector and no occlusion are selected to construct a candidate set of observation viewing angles for the region; Combining the candidate viewing angle set, the spatial distribution of the fruit, and the motion constraints of the collection device, the path of all effective observation postures is optimized to form a continuous, reachable, and unoccluded camera pose control sequence; According to the camera pose control sequence, the image collection task is executed in sequence, and the shooting position, posture angle, and image quality information are recorded after each group of shooting actions. When the image quality is not up to standard, a backup posture is called to retake the shot, and a complete multi-view image sequence is constructed.
[0009] Preferably, the steps of generating the maturity preliminary determination label are as follows: According to the shooting position and direction parameters recorded in the camera viewing angle control sequence, the fruit targets appearing in the multi-view images are identified and located, the spatial coordinates of the fruit center point are calculated using the camera parameters, and it is determined whether it is the same fruit according to the spatial position proximity and target size change rate. All corresponding images are grouped; A unique identity number is assigned to each group of images, and skin features such as color distribution, glossiness, texture density, shape contour, and patch area are extracted from all images under this number frame by frame to form a feature data table with viewing angle annotation; The multi-frame features under the same number are aggregated using standardized processing and weighted fusion, a cross-view consistency evaluation system is established based on the viewing angle direction weight, and finally the maturity preliminary determination label and the corresponding confidence value are generated; The data structure containing image path, shooting parameter, spatial coordinate, feature data, and maturity label is established with identity number as index, which runs through the stages of collection, identification, verification, and decision execution, realizing consistent tracking and calling of the whole process data.
[0010] Preferably, the steps of generating the maturity confidence interval are as follows: Extract image feature data of the same fruit under multiple perspectives based on the unique identity number, including color, luster, epidermis texture density, sugar spot area and contour shape; Numerical analysis is performed on the image features under all perspectives to calculate the maximum value, minimum value and fluctuation range respectively; According to the numerical fluctuation results, a maturity confidence interval set composed of minimum value and maximum value is constructed, each interval corresponds to a physical feature item and has clear numerical boundaries; Compare the fusion feature value in the preliminary judgment label with the maturity confidence interval. When all fusion values fall within the corresponding confidence interval, it is determined to be consistent, otherwise the re-observation strategy is triggered.
[0011] Preferably, when the stability of the maturity confidence interval does not reach the preset threshold, the re-observation strategy is executed as follows: Determine whether the fluctuation range of key features in the maturity confidence interval exceeds the set threshold. When the red channel gray value fluctuation is greater than 8 or the sugar spot area fluctuation exceeds 3% of the fruit skin area, perform a retake operation to control the image acquisition device to fine-tune the spatial position and pitch angle based on the current posture, and collect supplementary images to complete the key maturity feature area; Combine the three-dimensional spatial position and posture information of the fruit in the retake image to evaluate the current accessibility and clamping conditions of the gripper, including the distance between the clamping center and the fruit center, the clamping width, the clamping angle tolerance and the blocking interference in the clamping path, to determine whether the picking action can be executed; Include the retake image in the data set under the original identity number, update the maturity features and confidence intervals, and if all key features meet the stability requirements of the confidence intervals, combine the clamping state evaluation results to regenerate the picking action path and execution timing. If it still does not meet the requirements, the automatic picking is aborted and marked as abnormal.
[0012] Preferably, after picking is completed, the result feedback feature library is updated based on the difference between picking and prediction to dynamically update the feature library, realizing the closed-loop optimization of recognition control as follows: Collect actual result information during the picking stage, including whether the clamping is successfully completed, whether there is a slip or misgrasp phenomenon, whether the fruit is intact, whether the maturity state is consistent with the prediction result confirmed by artificial or sensing, and compare the information according to the fruit unique number and prediction features to identify the image perspective and corresponding feature item that has a significant residual error; Based on the image number and imaging parameters identified as having errors, the mapping area of the corresponding perspective in the fruit three-dimensional surface model is extracted, and the perspective sensitivity index value of the mapped area in the multi-angle spectral texture feature library is updated to reflect the state that this perspective has poor recognition stability in the current environment; The corrected sensitivity index is used to reconstruct the visibility prediction map, and the camera pose arrangement and feature fusion weight strategy in subsequent image acquisition are adjusted accordingly to ensure that the recognition results are more stable in subsequent tasks. The picking path and execution parameters are continuously optimized to build a feedback-driven closed-loop update mechanism.
[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention, through the construction of a multi-angle spectral texture feature library and viewpoint sensitivity modeling, enables the system to identify the visibility of key maturity features under different viewpoints. Combined with a camera viewpoint arrangement mechanism, it effectively avoids the problem of occlusion of mature areas. Furthermore, through fruit identity binding, multi-viewpoint consistency verification, and maturity confidence interval construction, it ensures the consistency and reliability of the identification judgment of the same fruit. When the identification is uncertain, a re-observation strategy and a real-time update mechanism for control commands are used to achieve coordinated adaptive adjustment of identification and action. Finally, the harvesting results feed back into the feature library, constructing a closed-loop optimization mechanism that can continuously evolve with the operation results. The overall solution not only solves the problems of misjudgment, oscillation, and low efficiency caused by viewpoint occlusion in the existing technology for fruit maturity identification, but also significantly improves the system's environmental adaptability, identification robustness, and operation execution accuracy. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0015] Figure 1 This is a flowchart of the automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to the present invention. Detailed Implementation
[0016] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0017] This invention provides, for example Figure 1 The illustrated method for automated harvesting and grading of fruits and vegetables based on AI vision and dynamic control includes the following steps: A multi-angle spectral texture feature library is constructed. In the multi-angle spectral texture feature library, the view sensitivity index is calculated for the surface image features corresponding to different viewing angles of the fruit. Based on the view sensitivity index, a visibility prediction map is generated to mark the ripe areas on the fruit surface that are easily obscured at different viewing angles. To achieve stable identification of fruit ripening areas under multi-angle conditions and avoid the loss of feature information due to changes in viewing angle or occlusion, a method is proposed to construct a multi-angle spectral texture feature library and generate a viewing angle sensitivity index and visibility prediction map, serving as crucial perceptual support in the automated fruit and vegetable harvesting process. The specific implementation steps are as follows: Multi-angle image acquisition was conducted to obtain the visual appearance of the target fruit from different observation angles. The specific steps were as follows: In the orchard or experimental platform, a set of motorized turntables with rotation and pitch capabilities were set up, and multiple standard varieties of fruit were fixed at the center of the turntables to ensure the fruit's position remained constant during the acquisition process. Using a combination of a color camera and a near-infrared camera with at least 12 megapixel resolution, image data was acquired around the spherical surface of the fruit in no fewer than twelve directions under fixed lighting conditions. Each direction corresponded to a specific combination of azimuth and pitch angles, such as 0°, 30°, 60°, 90°, 120°, 150°, and 180°. Color and near-infrared images were acquired at each angle to obtain information on the peel color and internal component reflectance characteristics (such as moisture distribution and differences in reflectance in saccharified areas), respectively. Before each acquisition, a uniform camera focal length (e.g., 35 mm) and a fixed exposure time (e.g., 1 / 100 second) were set to ensure the consistency and comparability of the images. After all images are acquired, they are classified and stored according to fruit number, viewpoint number, and spectral type to form a multi-angle raw image data set.
[0018] Image feature extraction and viewpoint observability analysis were performed. Color images from each viewpoint were input into an image processing program to extract the main color components of the fruit peel, including the spatial distribution of red, green, and blue components. Edge detection technology was used to identify patches, color difference regions, and texture change boundaries. In near-infrared images, the reflectance intensity distribution of the peel at specific wavelengths (e.g., 850 nm and 940 nm) was extracted to estimate the ripeness differences at different parts of the fruit. The color feature regions in each image were correlated with their spatial locations and mapped onto a three-dimensional fruit surface model. At each point on the fruit surface, it was recorded whether it was effectively observed from all viewpoints and whether it contained key ripeness features (e.g., color spots, sugaring points, uniform color areas). This information was statistically analyzed to determine the visibility frequency and feature appearance frequency of each surface region across all viewpoints. Furthermore, the visibility frequency of each surface region was divided by the total number of viewpoints to obtain its observable proportion. Combined with feature changes, a viewpoint sensitivity index was calculated to characterize the degree of change in observed features in a given area due to angle variations. The sensitivity index is represented by a continuous value from 0 to 1. The higher the value, the more easily the area is affected by the viewing angle and thus distorted or obscured.
[0019] Based on the calculated perspective sensitivity index, a visibility prediction map of the fruit surface is constructed. First, the three-dimensional fruit surface is mapped onto a two-dimensional image plane using polar coordinate projection, unfolding the fruit sphere into a two-dimensional map in latitude and longitude. Each image pixel represents a specific physical region on the fruit surface, corresponding to a unique spatial coordinate. The previously calculated perspective sensitivity index is then entered into this coordinate system, with different gray levels used to represent the index's intensity. Further, regions with a sensitivity index higher than 0.7 are selected for boundary tracking processing, enclosing them as continuous regions. Edges are then drawn using polygon boundary fitting technology, forming a visual representation of high-occlusion-risk areas. To ensure compatibility with practical applications, the image also labels the approximate direction of each region from the actual picking perspective, used for precise avoidance of occluded areas in subsequent posture planning. The final visibility prediction map is stored at a standard resolution (e.g., 640x640 pixels) and bound to the number of each fruit, serving as a spatial reference map for the fruit's multi-angle observability.
[0020] Based on the distribution of sensitive areas marked in the visibility prediction map, a visibility safety boundary for the fruit surface is established, and a data acquisition strategy for this boundary is constructed. In implementation, for any mature feature area (e.g., a distinctly red patch), it is required to be imaged from at least three low-sensitivity indices (less than 0.3) during the acquisition process to ensure stable acquisition of its maturity features. During camera position planning, the perspective corresponding to the high-visibility area marked in the visibility prediction map is prioritized. Simultaneously, the shooting angle is further fine-tuned based on the distribution of obstructions around the fruit (e.g., leaf orientation, pedicel position) to reduce potential obstructions. After this acquisition strategy is executed, the perspective parameters of each successfully observed mature area are automatically recorded, and the system continuously compares whether the visibility redundancy requirement is met (i.e., multiple angles confirm the maturity features of the same area). If a feature area does not reach the preset observation redundancy number, the system arranges for the camera to observe the area again from different directions until the perspective redundancy requirement is met, ensuring that subsequent identification and judgment are based on sufficient and stable information.
[0021] The main function of this step is to establish the mapping relationship between the spectral and texture features of the fruit under multi-view conditions, and based on this, construct a quantifiable view sensitivity index and visibility prediction map, thus providing a robust perceptual foundation for subsequent visual recognition and control planning. In traditional fruit recognition, the system often relies on a single-view image to determine maturity, which is easily affected by factors such as asymmetrical distribution of fruit surface features and occlusion by leaves and pedicels, resulting in the inability to continuously observe key maturity features and causing recognition errors. This step, however, collects image data of the fruit from multiple directions, extracts features, and models sensitivity, enabling the quantification of the visibility and occlusion risk of each surface area. By generating a visibility prediction map, it accurately identifies which areas are most likely to be lost under which viewpoints, fundamentally solving the blind spot problem of traditional methods in response to viewpoint changes. This step not only improves the targeting and efficiency of subsequent image acquisition strategies but also provides higher data robustness and recognition decision reliability for the entire automated harvesting system, making it a key preliminary step for achieving stable fruit recognition in dynamic, multi-view environments.
[0022] Camera viewpoint arrangement is implemented based on visibility prediction maps. By formulating camera pose control sequences, mature areas with high viewpoint sensitivity index are kept visible with minimal occlusion during the acquisition process. To ensure that the ripe fruit area remains visible with minimal occlusion during harvesting image acquisition, targeted camera viewpoint choreography is required based on the previously generated visibility prediction map. This process ensures that the image acquisition process covers all high-viewpoint-sensitive areas of the fruit surface as completely as possible by determining the camera's movement path and shooting posture sequence in space. The specific implementation method for this step is as follows: Based on the visibility prediction map constructed in the previous stage, regions on the fruit surface with high sensitivity to ripening characteristics were identified. Specifically, all pixels with a grayscale value greater than or equal to 180 were extracted from the 2D visibility prediction map; these pixels corresponded to fruit surface regions with a viewpoint sensitivity index higher than 0.7. In the 3D spherical reconstruction model corresponding to the same fruit number, the mapping relationship between pixel position and 3D coordinates was used to accurately locate these highly sensitive pixels in the spherical coordinate system of the fruit surface. For each highly sensitive region, its spatial coordinates (including spherical radius, azimuth, and pitch angles), surface normal vector direction, and the average visibility value of adjacent surface patches were recorded. These parameters comprehensively describe the observable characteristics and occlusion trends of the region under different viewpoints.
[0023] Based on the positional relationships and surface orientation of highly sensitive areas, a set of candidate camera viewing angles suitable for capturing images of these areas is selected. The specific process is as follows: using the surface normal vector of each highly sensitive area as a reference direction, a conical observation space is constructed in the opposite direction with a step angle of 30°. Several preset camera position points are set in this space, each representing a specific shooting posture, including its position coordinates and orientation direction. At each candidate observation point, the angle between the line of sight at that point and the surface normal vector of the highly sensitive area is calculated. An angle less than 45° and without obstructions (judged by analyzing the sensitivity index values of adjacent areas in the visibility prediction map around the fruit surface) is considered a valid observation posture. In this way, at least three valid shooting postures are generated for each highly sensitive area, forming a candidate set of viewing angles.
[0024] By combining the candidate set of observation viewpoints, the overall spatial distribution of the fruit, and the mechanical movement capabilities of the acquisition mechanism, path optimization is performed on all effective shooting postures to generate a continuous, reachable, and unobstructed camera pose control sequence. Specifically, all observation postures are sequentially connected in three-dimensional coordinate space. The required movement distance, rotation angle, and potential occlusion risk (assessed based on the positional data of other fruits, leaves, and branches around the fruit) required for the camera to move from its current position to the next target posture are calculated. Under the conditions of satisfying the allowable rotation angle range of the mechanical structure (e.g., pitch ±60°, horizontal rotation ±90°) and movement speed limits (e.g., movement not exceeding 15 cm per second), the posture execution order is arranged according to the principle of shortest path distance and lowest occlusion risk, forming the final shooting path. Each posture node is assigned a unique number and estimated arrival time in the path, and the corresponding target observation area number is recorded to ensure that the path planning has a clear execution objective and time constraints.
[0025] The generated camera pose control sequence is converted into specific image acquisition actions. The control device sequentially drives the acquisition equipment to designated positions based on the target points in the control sequence, adjusting the camera pose to the desired orientation. Upon reaching each pose point, an image is immediately acquired, recording the current camera's spatial position information, shooting angle, captured area number, and image quality indicators (such as sharpness, average brightness, and occlusion rate). If the image quality does not meet the preset standards (e.g., sharpness below a set threshold or occlusion rate exceeding 15%), a reshoot mechanism is initiated: a backup observation pose for the same area is called, and the next priority shooting angle is selected from the candidate set to re-acquire the image. All acquired images correspond one-to-one with the fruit number, high-sensitivity area number, and viewpoint number, constructing a complete multi-view image sequence for subsequent recognition, fusion, and maturity assessment.
[0026] The purpose of this step is to proactively orchestrate the camera's shooting angles and movement paths in space, based on the previously constructed visibility prediction map. This ensures that ripe areas on the fruit surface, which are highly sensitive to visual changes and easily obstructed, are fully observed with minimal occlusion during image acquisition. Traditional techniques typically use fixed or random viewing angles to photograph the fruit, failing to dynamically adjust shooting strategies according to changes in the visibility of fruit features at different angles. This leads to the omission of crucial ripening information due to occlusion or viewing angle deviations, resulting in identification errors and misjudgments during harvesting. This step extracts the three-dimensional spatial positions of highly sensitive areas on the fruit surface, selects multiple observation postures with optimal visibility for these areas, and optimizes the camera movement path and acquisition sequence by considering the fruit's spatial distribution, the layout of surrounding obstructions, and the robotic arm's capabilities. The resulting camera pose control sequence not only ensures redundant observation of ripening features but also effectively avoids the effects of occlusion, making the acquisition process highly adaptable and controllable. This is a crucial supporting element for ensuring the stability of fruit identification, the accuracy of action decisions, and the overall robustness of the system.
[0027] After completing the camera perspective arrangement, a spatiotemporal binding mechanism for fruits is established, a unique identifier is assigned to each fruit to be identified, and the features of fruit images collected from multiple perspectives are aggregated under the unique identifier to generate a preliminary maturity judgment label. After completing camera viewpoint arrangement and image acquisition, the acquired multi-view image data needs to be uniformly identified and classified as fruit targets, and a spatiotemporal binding mechanism for fruits needs to be constructed based on this. By assigning a unique identifier to each fruit to be identified and aggregating all image features corresponding to that identifier, a stable judgment of the fruit's ripeness status can be achieved. This process is a prerequisite for subsequent consistency verification, confidence fusion, and action execution. The specific implementation method is as follows: Based on the shooting position and orientation parameters recorded by the camera view control sequence, the same physical fruit individual is identified and located from multi-view images. Specifically, the process involves: taking the identified fruit region in each frame as the target, recording its two-dimensional position in the image plane (i.e., the coordinates of the center point in the image coordinate system), edge contour shape (extracting the circumscribed ellipse through edge detection), and target size (calculating the pixel span in the horizontal and vertical directions). Using the camera's three-dimensional position coordinates, orientation angle, and camera intrinsic parameters (focal length, principal point position) during acquisition, the three-dimensional position of the fruit's center point in the world coordinate system is calculated inversely. This method yields the observation coordinates of the fruit in three-dimensional space across all images. If the spatial distance between observation points in different images is less than a set threshold (e.g., distance difference less than 50 mm), and the corresponding target size change rate does not exceed 20%, they are determined to be the same fruit. Images are sequentially matched across all view images, grouping those meeting the criteria, with each group corresponding to an independent fruit entity. Each group is then assigned a unique identification number, such as F001, F002, F003, etc., for feature classification and data tracking in subsequent steps.
[0028] After assigning identification numbers, feature extraction is performed on the image group corresponding to each identification number. In each frame, the following physical and appearance features are extracted from the target area of the fruit: 1) Color distribution features, including the mean and standard deviation of the red, green, and blue channels, used to describe the overall hue and ripening trend of the peel; 2) Gloss information, using the pixel distribution and grayscale mean of the bright areas in the image to determine the smoothness and reflectivity of the fruit peel; 3) Texture density, analyzing the uniformity of the peel surface by detecting the frequency of edge changes and the density of spot distribution; 4) Shape features, assessing the integrity and degree of deformation of the fruit by calculating the ratio of the major and minor axes of the circumscribed ellipse of the fruit edge; 5) Patch features, identifying the presence of typical ripening signs such as localized sugaring, red spots, and browning through regional color clustering and shape contour extraction. All feature data for each frame are archived sequentially, and their corresponding viewpoint parameters, camera position, and shooting time are labeled to establish a complete image feature record table.
[0029] The image features of multiple frames of the same fruit target are aggregated to form a maturity judgment criterion with cross-view consistency. During the aggregation process, each type of feature is first standardized. For example, color channel values are normalized to the range of 0 to 1, gloss grayscale values are adjusted to a unified brightness reference frame, texture density is expressed as the number of edge jumps per 100 pixels, and patch area is expressed as a percentage of the fruit's surface area. Then, a weighted fusion is performed based on the angle between the shooting viewpoint and the target surface orientation. For example, the weight for frontal observation is set to 1.0, the weight for oblique side observation is set to 0.6, and the weight for back view is set to 0.3, ensuring that high-confidence viewpoints play a dominant role in the aggregation result. After comprehensive evaluation of the fused data, a preliminary maturity judgment label for the fruit is generated. This label can adopt a three-level classification method: "immature," "critically mature," and "fully mature," and includes a maturity confidence value, expressed as a percentage of recognition accuracy (e.g., 93%). This label information is permanently bound to the identification number, forming a preliminary identification result set.
[0030] A sustainable data structure is constructed, using the identification number as an index to categorize and store image paths, camera parameters, spatial locations, feature data, and maturity labels in key-value pairs. This structure not only supports rapid retrieval in subsequent consistency verification steps but also allows for triggering a re-observation strategy to call back the original image group for repeated identification when the system determines that the maturity label has low credibility. Furthermore, this data structure allows for state backtracking after the fruit is harvested, used for identification model correction and misjudgment analysis. The identification number, as the core index of the entire data flow, runs through the entire process of acquisition, identification, judgment, and execution, and is a key mechanism to ensure the high reliability and controllability of the entire system.
[0031] This step aims to establish a spatiotemporal binding mechanism for fruits, enabling accurate identification and feature aggregation of the same fruit target in multi-view images. Based on this, a preliminary maturity assessment label is generated, providing a foundation for subsequent consistency verification and harvesting decisions. In multi-view image acquisition, the appearance characteristics of fruits change significantly from different angles and may be partially occluded. Failure to accurately identify and uniformly classify fruits can lead to misidentification of the same fruit as multiple targets, causing maturity judgment conflicts or the risk of repeated harvesting. This step assigns a unique identification number to each fruit through spatial location matching and contour feature association, and structurally summarizes its image data from multiple perspectives. Simultaneously, multi-dimensional visual features, including color, gloss, texture density, shape contour, and patch regions, are extracted and weighted according to the viewing angle to ensure that the fused information represents the overall true state of the fruit. The resulting preliminary maturity label not only has higher stability and accuracy but also provides the necessary data foundation for improving the confidence of subsequent identification results and closed-loop control, making it a key link in constructing the core perception logic of the intelligent harvesting system.
[0032] Based on the initial maturity assessment label, a multi-view consistency check is performed. Image features from multiple views are fused to generate a maturity confidence interval, which is used to eliminate recognition errors caused by single-view acquisition. To ensure the stability of fruit maturity labels across different viewing angles and effectively avoid identification errors caused by partial occlusion, lighting differences, or imaging deviations, after generating the initial maturity determination label, the consistency of feature data acquired from the same fruit at multiple viewpoints needs to be checked, generating maturity confidence intervals for assessing the stability of the identification results. This process confirms the reliability of the maturity labels by comparing and analyzing the fluctuation range and common trends of various features in images from different viewpoints. The specific implementation method is as follows: Based on the fruit IDs whose identities have been bound in the previous step, image feature data for the fruit is extracted from all valid viewpoints. Each image feature consists of five parts: image ID, shooting angle information, three-dimensional position of the fruit's center point, color features, gloss features, skin texture density, distribution and area of saccharified regions, and fruit outline shape. Taking color features as an example, the average pixel value and standard deviation of the red, green, and blue channels are recorded; taking gloss features as an example, the maximum brightness value and area of the bright areas in the image are statistically analyzed; taking saccharified regions as an example, the proportion of the saccharified area to the entire peel area is calculated and its specific location in the image is recorded. All of the above features are described by manually defined measurable units and do not involve abstract function forms, ensuring that each feature has consistent comparability across different images.
[0033] Data on similar features from various perspectives were compiled and organized to establish a unified multi-view feature set. Using fruit number as an index, the data for color, gloss, texture, and sugar-coated area corresponding to all images under that number were aggregated into a single table, arranged horizontally with perspective number as rows and feature item as columns. After data organization, the maximum, minimum, and average values of each feature were calculated to obtain its fluctuation range. For example, if the average red channel pixel values of a certain fruit under seven perspectives are 172, 174, 173, 169, 175, 171, and 170 respectively, then the maximum value is 175, the minimum value is 169, the average value is 172, the fluctuation range is 6, and the fluctuation ratio is 3.5%. This fluctuation calculation was performed once for each feature to comprehensively assess the stability of various maturity features.
[0034] Based on the collected multi-view feature fluctuations, a maturity confidence interval for the fruit is constructed. This confidence interval is not based on a probability distribution model, but rather directly defined based on the range of physical quantity values. For example, if the red channel pixel value fluctuation range is within a set 5% threshold, the confidence interval is set to [minimum, maximum], such as [169, 175]. If the sugar spot area is concentrated between 12% and 14% of the peel area across all views, the confidence interval for the sugary area is defined as [12%, 14%]. This interval setting is applied to each feature, ultimately forming a set of quantifiable confidence intervals. Each item in the set is represented by two specific boundary values and has direct physical meaning. This set of confidence intervals is then compared with the fused feature values used for initial label generation. If the fused value falls within the middle region of all confidence intervals, and no single feature deviates from the confidence range, the maturity determination is considered consistent. Conversely, if the fused value of a key feature exceeds the confidence interval boundary, it is considered that there is a risk of inconsistency in identification.
[0035] Based on the confidence interval comparison results, a consistency judgment conclusion is output and associated with the fruit number. If all feature items are within the confidence interval, it indicates that the recognition features under different viewpoints are highly stable, the original preliminary judgment label is confirmed as a valid result, and can be used to guide the execution order and parameter configuration of subsequent harvesting operations. If a feature item deviates significantly from the confidence interval, such as gloss being more than 50% higher in one viewpoint than in other viewpoints, or the area of sugar spots deviating significantly from other image values in a single frame image exceeding a set threshold, this abnormal feature is automatically recorded, and the fruit is marked as having an unstable recognition state, simultaneously entering the next step of the re-observation mechanism. During this process, the original image data is not modified, nor are abstract reasoning methods used; instead, quantitative data is used as the basis to ensure that the judgment criteria are verifiable and physically repeatable.
[0036] The purpose of this step is to construct a maturity confidence interval through multi-view consistency verification, thereby improving the stability and reliability of fruit maturity judgment results and solving misjudgment problems caused by information bias, local occlusion, or lighting differences due to single-view acquisition. In the previous steps, although preliminary maturity judgment labels have been generated based on multi-view images, there may still be cases where individual viewpoint features are abnormal or data deviates from reality. To avoid such incidental factors affecting system decisions, this step performs a horizontal comparison and fusion of multiple physical features extracted from the same fruit under different views, such as color, gloss, texture, and patch area. By setting a quantifiable fluctuation range, a confidence interval with physical boundaries is constructed, and the stability and concentration of feature data are determined accordingly. If all feature fusion values are within the confidence interval, the judgment is consistent and the label is reliable; otherwise, it indicates a deviation in identification, requiring further observation or correction. This mechanism ensures that fruit identification results are not dominated by incidental images from a single angle, but are generated on the basis of global consistency, significantly enhancing the robustness of the system's judgment in complex orchard environments. It is a key supporting link for achieving highly reliable automated harvesting decisions.
[0037] When the stability of the maturity confidence interval does not reach the preset threshold, a re-observation strategy is executed. By fine-tuning the camera pose and combining the calculation results of the current gripping operation window of the robotic arm, the control decisions and action timing of the picking are updated. When fruit maturity identification still exhibits instability after multi-view feature fusion—meaning the confidence interval width exceeds a preset threshold or key features show significant conflicts across different views—it indicates uncertainty in the current identification result. In this situation, to avoid erroneous picking due to identification bias, a re-observation strategy must be initiated immediately. The re-observation strategy includes fine-tuning the camera pose and jointly evaluating the current state of the robotic arm's gripping window, thereby updating the control decisions and execution timing of the picking action. The specific implementation method is as follows: The system determines whether the confidence interval of the current fruit maturity recognition result meets the stability requirements. Taking the fluctuation of grayscale values in the color channel as an example, if the average pixel value fluctuation range of the red channel exceeds a set threshold (e.g., 8 grayscale units) across all views, or if the area ratio of the saccharified region differs by more than 3% of the total fruit surface area across different views, the confidence interval is considered unstable. After triggering the re-observation strategy, the missing information areas in the image to be supplemented are identified. By comparing the edge contours, surface orientation, and distribution of highly sensitive areas of the fruit in images taken from different views, areas that are occluded or not fully observed are identified. Subsequently, based on the current spatial position and orientation angle of the camera, its attitude parameters are fine-tuned. Specific fine-tuning operations include: shifting the camera position horizontally towards the target fruit by no more than 20 mm, adjusting the vertical height by no more than 15 mm, and adjusting the pitch angle within ±10 degrees to avoid excessive displacement causing occlusion or shooting outside the field of view. The adjusted camera posture should ensure that the target area is centered in the image and maintain sufficient image resolution. Once the image quality indicators (such as focal length, brightness, and contrast) meet the acquisition standards, a reshoot operation should be performed to acquire supplementary images and record the shooting parameters, shooting time, and viewpoint number of the images.
[0038] After image re-acquisition, the gripping status of the robotic arm needs to be incorporated into the comprehensive judgment of the re-observation strategy. First, the three-dimensional spatial position and attitude parameters of the fruit in the re-acquisition image are extracted. The spatial distance between the fruit's centroid and the gripper's center is calculated. If this distance is less than the set effective gripping radius (e.g., 30 mm), it is considered reachable. Next, the major and minor axis lengths and tilt angles of the fruit's circumscribed contour are extracted. Combined with the current gripper's maximum opening width (e.g., 80 mm) and the allowable gripping surface tilt angle (e.g., 20 degrees), it is determined whether the fruit is within the grippable posture range. If the spatial position and attitude conditions are met, a gripping window evaluation is performed. Specific evaluation content also includes the positional interference of obstacles around the fruit (e.g., leaves, pedicels): by comparing the edge distribution density and color distribution in the image, it is determined whether there are interfering objects overlapping within the gripping path; if the area obscured by the interfering object is less than 15% of the gripping area, it is still considered grippable. After the evaluation is passed, the re-acquisition image is incorporated into the dataset under the original identification number, maturity features are re-extracted, and the confidence interval is jointly updated with the original feature data.
[0039] After the re-captured images and clamping window data are updated, the picking action control decision is updated again. First, it is confirmed that the updated confidence interval meets stability conditions for all key maturity features. For example, if the fluctuation range of the red channel grayscale value is reduced to within 6, and the fluctuation rate of the sugar spot area is reduced to within 2%, the recognition result is considered reliable. Under the premise that the maturity is stable and the clamping window is effective, the fruit is re-added to the list of targets to be picked, and a new picking priority number is assigned. The priority is determined by the fruit maturity level, recognition confidence score, distance from the current position of the robotic arm, and the density of surrounding fruits, ensuring the shortest robotic arm movement path and minimum energy consumption. The control device then reorders the picking path, sets control parameters such as the gripper contact point position, opening and closing speed, and clamping force, and adjusts the action time sequence of the robotic arm joints, including the start time, movement time, clamping duration, and reset time, to ensure that the picking process completely corresponds to the new recognition data. If the recognition results remain unstable after re-shooting, or the clamping posture continues to fail to meet requirements, the automatic harvesting process will be terminated, the fruit will be marked as an abnormal target, and manual review will be required. Simultaneously, the abnormal mark will be stored in the recognition database for subsequent recognition model optimization.
[0040] The purpose of this step is to improve the accuracy of fruit maturity identification and the safety of harvesting decisions when the results are unstable, i.e., the feature data differs significantly from multiple perspectives, making it impossible to form a reliable maturity confidence interval. This is achieved by implementing a re-observation strategy. First, the step involves fine-tuning the camera's spatial position and angle to supplement images of key areas, acquiring high-quality images of obscured or poorly imaged areas. Next, the step considers the current gripping operation window of the robotic arm to assess the fruit's spatial position, surface posture, and the accessibility of the gripping path, ensuring the fruit remains graspable after the supplementary images. Finally, based on the feature data extracted from the new images, the maturity confidence interval is reconstructed, and the control commands and execution sequence of the harvesting action are updated accordingly, ensuring the action path, gripping angle, and identification results match. Unlike traditional technologies that directly abandon or blindly harvest due to unclear judgments, this step achieves closed-loop control through image supplementation and action optimization. This not only avoids misharvesting but also improves the overall identification system's fault tolerance and dynamic adjustment capabilities in complex orchard environments, making it a crucial element in ensuring system robustness and intelligence.
[0041] After the picking action is completed, the picking results are fed back to the multi-angle spectral texture feature library. The perspective sensitivity index is corrected according to the residual between the picking results and the prediction results, forming a feature library dynamic update mechanism driven by the picking effect, and constructing a closed-loop optimization process for recognition and control. After the fruit is harvested, relying solely on the image recognition prediction results from before harvesting is insufficient to continuously improve the system's recognition accuracy and control stability. To achieve self-correction and continuous optimization of recognition and control capabilities, it is necessary to compare and analyze the actual harvesting results with the predicted results. This feedback is then used to refine the sensitivity description indicators of the fruit surface region under different viewing angles, ultimately constructing a closed-loop optimization mechanism for recognition and control driven by harvesting feedback. The specific steps are as follows: After the fruit is harvested, multiple feedback results from the harvesting stage are collected and compared one-to-one with the prediction results output by the image recognition process before harvesting. Harvesting feedback results include: whether the fruit was successfully grasped and harvested; whether there was any displacement, slippage, or accidental grasping of non-target objects during the grasping process; whether the fruit is intact after harvesting; whether the fruit is ripe; and confirmation of the ripeness level by manual intervention or high-precision sensors. Prediction results include: ripeness labels, fused image feature values, color channel values extracted from images from various viewpoints, texture features, sugar spot area distribution, and peel gloss. All this information is indexed by the fruit's unique ID for result archiving and comparison. For example, if the predicted sugar spot area is 15% of the fruit's surface area during pre-harvest recognition, but the actual measured area after harvest is 6%, it indicates an overestimation of feature extraction, with viewpoint occlusion or feature recognition anomalies. Through this residual analysis, it is possible to accurately identify which features had errors during the recognition stage and trace them back to their respective viewpoints and image IDs.
[0042] To address the aforementioned sources of recognition error, the image perspective with the most significant deviation from the prediction is identified. Based on the imaging range corresponding to this perspective on the fruit surface, the perspective sensitivity index of the corresponding region in the multi-angle spectral texture feature library is corrected. The specific process is as follows: Based on the camera spatial pose recorded during image acquisition, the shooting perspective number corresponding to each image is assigned, and the image imaging area is mapped onto the three-dimensional surface model of the fruit to determine its coverage area. For example, for fruit number F017, if the image perspective number V03 covers the top area of the fruit, and this image outputs an incorrect feature value in sugar spot recognition, and post-acquisition feedback confirms that the image quality at this perspective exhibits high reflectivity, occlusion, or blurring, then the perspective sensitivity index of this region under perspective V03 is increased from the original value (e.g., 0.45) to a new value (e.g., 0.75) to reflect that this perspective is more likely to cause recognition errors in the current environment. The corrected sensitivity index directly affects the perspective arrangement strategy in subsequent image acquisition stages. In the next round of shooting, the system will prioritize collecting information from perspectives with lower sensitivity and higher recognition stability, or arrange reshooting postures for the highly sensitive perspective to improve recognition redundancy.
[0043] The corrected viewpoint sensitivity index is re-input into the multi-angle spectral texture feature library to update the feature library content. Simultaneously, a new visibility prediction map is constructed based on the latest sensitivity index. This prediction map uses the three-dimensional spherical surface of the fruit as a projection, visualizing the sensitivity index of each fruit surface region as a grayscale difference. Higher grayscale values indicate greater sensitivity to viewpoint changes and lower recognition reliability. Subsequently, the camera viewpoint arrangement order is regenerated, prioritizing the acquisition of main viewpoint images from low-sensitivity areas, and implementing auxiliary acquisition compensation mechanisms for high-sensitivity areas, such as refining the shooting angle, fine-tuning the focal length, or adjusting the light source position. Furthermore, in the next round of image feature fusion judgment, image features identified as having high-sensitivity viewpoints are assigned lower fusion weights, making the overall judgment more reliant on stable viewpoints and reducing the impact of occasional deviations on the final label output. Through this continuous closed-loop process, after each picking action, the system can automatically optimize the recognition process, acquisition parameters, and control strategies in real time based on the picking effect feedback, thereby continuously improving overall recognition accuracy and action execution success rate.
[0044] The purpose of this step is to feed back the actual results of fruit harvesting to a multi-angle spectral texture feature library. Based on the differences between the actual harvesting results and the initial image recognition predictions, the system dynamically adjusts the viewpoint sensitivity index of the fruit surface region from various perspectives, thereby continuously optimizing the decision-making basis for recognition and control, and constructing a closed-loop optimization mechanism driven by harvesting results. By comparing and analyzing actual feedback data such as whether the harvest was successful, whether the clamping was accurate, and whether the maturity matched the target, with the maturity labels, feature values, and image quality output from the recognition stage, the system can pinpoint the viewpoints from which recognition errors originate, identifying areas of recognition deviation caused by occlusion, light spots, or blurring in specific images. The system then adjusts the viewpoint sensitivity index of these areas in the feature library, increasing the recognition weight for their uncertainty, and avoiding reliance on incorrect viewpoints in subsequent recognitions. Simultaneously, by reconstructing the visibility prediction map and optimizing the camera acquisition path, the system further improves the information acquisition quality and motion control accuracy in future tasks, achieving collaborative adaptive evolution between the recognition process and harvesting control, making the entire automated process more learning-capable and long-term stable.
[0045] The aforementioned AI-based vision and dynamic control-based automatic fruit and vegetable harvesting and grading method transforms fruit maturity recognition from a single-view static judgment to a multi-view dynamic fusion, significantly improving the accuracy and stability of recognition results in complex orchard environments. This method, through the construction of a multi-angle spectral texture feature library and viewpoint sensitivity modeling, enables the system to identify the visibility of key maturity features from different perspectives. Combined with a camera viewpoint arrangement mechanism, it effectively avoids the problem of occlusion of mature areas. Furthermore, through fruit identity binding, multi-view consistency verification, and maturity confidence interval construction, it ensures the consistency and reliability of the recognition judgment for the same fruit. When recognition is uncertain, a re-observation strategy and a real-time update mechanism for control commands enable coordinated adaptive adjustment of recognition and action. Finally, the harvesting results feed back into the feature library, constructing a closed-loop optimization mechanism that can continuously evolve with the operational results. The overall solution not only solves the problems of misjudgment, oscillation, and low efficiency caused by viewpoint occlusion in existing technologies for fruit maturity recognition, but also significantly improves the system's environmental adaptability, recognition robustness, and operational accuracy.
[0046] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A method for automatic harvesting and grading of fruits and vegetables based on AI vision and dynamic control, characterized in that, Includes the following steps: A multi-angle spectral texture feature library is constructed. In the multi-angle spectral texture feature library, the view sensitivity index is calculated for the surface image features corresponding to different viewing angles of the fruit. Based on the view sensitivity index, a visibility prediction map is generated to mark the ripe areas on the fruit surface that are easily obscured at different viewing angles. Camera viewpoint arrangement is implemented based on visibility prediction maps. By formulating camera pose control sequences, mature areas with high viewpoint sensitivity index are kept visible with minimal occlusion during the acquisition process. After completing the camera perspective arrangement, a spatiotemporal binding mechanism for fruits is established, a unique identifier is assigned to each fruit to be identified, and the features of fruit images collected from multiple perspectives are aggregated under the unique identifier to generate a preliminary maturity judgment label. Based on the initial maturity assessment label, a multi-view consistency check is performed. Image features from multiple views are fused to generate a maturity confidence interval, which is used to eliminate recognition errors caused by single-view acquisition. When the stability of the maturity confidence interval does not reach the preset threshold, a re-observation strategy is executed. By fine-tuning the camera pose and combining the calculation results of the current gripping operation window of the robotic arm, the control decisions and action timing of the picking are updated. After the harvesting action is completed, the harvesting results are fed back to the multi-angle spectral texture feature library. The perspective sensitivity index is corrected based on the residual between the harvesting results and the prediction results, forming a dynamic update mechanism for the feature library driven by the harvesting effect, and constructing a closed-loop optimization process for recognition and control.
2. The automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to claim 1, characterized in that, The steps for generating the visibility prediction map are as follows: Perform multi-angle image acquisition operation. Using a color camera and a near-infrared camera mounted on an electric turntable with rotation and pitch functions, under constant light intensity, collect image data from no less than twelve angles around the spherical surface of the fruit at a fixed position. Each image is labeled with the corresponding azimuth and pitch angle information. The color distribution information, color spot boundary features, and reflectance intensity distribution in the near-infrared band extracted from each image are respectively mapped to specific physical regions in the three-dimensional surface model of the fruit. Based on this, the frequency of effective observation of each fruit surface region under various perspectives and the number of times the corresponding key ripening features are observed during the observation process are counted. Based on the observation frequency and the degree of feature variation, the view sensitivity index of each surface region is calculated, and a two-dimensional visibility prediction map is generated by polar coordinate expansion. Regions with view sensitivity indices higher than 0.7 are extracted with closed boundaries and labeled with spatial orientation. Based on the visibility prediction map, a visibility safety boundary for the fruit surface is constructed, and a multi-angle acquisition redundancy strategy is formulated to ensure that each mature area is stably imaged at least three low-sensitivity viewpoints, meeting the redundancy requirements for subsequent identification.
3. The automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to claim 2, characterized in that, The steps for implementing camera viewpoint arrangement based on visibility prediction maps are as follows: Based on the generated visibility prediction map, image regions with a view sensitivity index higher than 0.7 are extracted, and the corresponding pixel positions are mapped to the three-dimensional spherical model of the fruit to obtain the spatial coordinates and surface normal vector of each highly sensitive region. A conical observation space is constructed by reversing the surface normal vector of each highly sensitive area. Multiple preset shooting postures are set, and effective observation postures with an angle of less than 45 degrees to the surface normal vector and no obstruction are selected to construct a candidate set of observation viewpoints for that area. By combining the candidate viewpoint set, the spatial distribution of fruits, and the motion constraints of the acquisition device, the path of all effective observation postures is optimized to form a continuous, reachable, and unobstructed camera pose control sequence. Image acquisition tasks are executed sequentially according to the camera pose control sequence. After each set of shooting actions is completed, the shooting position, pose angle and image quality information are recorded. If the image quality is not up to standard, a backup pose is called for reshooting to construct a complete multi-view image sequence.
4. The automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to claim 3, characterized in that, The steps for generating a preliminary maturity assessment label are as follows: Based on the shooting position and direction parameters recorded in the camera view control sequence, the fruit targets appearing in the multi-view images are identified and located. The spatial coordinates of the center point of the fruit are calculated using the camera parameters. Based on the spatial proximity and the rate of change of the target size, it is determined whether they are the same fruit, and all corresponding images are grouped. Each group of images is assigned a unique identification number, and epidermal features such as color distribution, gloss, texture density, shape contour and patch area are extracted frame by frame from all images under that number to form a feature data table with viewpoint annotation; The standardization and weighted fusion methods are used to aggregate features from multiple frames under the same ID. A cross-view consistency evaluation system is established based on the view orientation weight, and finally, a preliminary maturity judgment label and corresponding confidence value are generated. A data structure containing image path, shooting parameters, spatial coordinates, feature data and maturity label is established using the identity number as an index. This structure runs through all stages of collection, identification, verification and decision execution, enabling consistent tracking and retrieval of data throughout the entire process.
5. The automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to claim 4, characterized in that, The steps for generating maturity confidence intervals are as follows: Based on the unique identification number, image feature data of the same fruit is extracted from multiple perspectives. The image feature data includes color, gloss, epidermal texture density, sugar spot area and outline shape. Numerical analysis of image features is performed from all perspectives, and the maximum, minimum and fluctuation ranges are calculated respectively. Based on the numerical fluctuation results, a set of maturity confidence intervals consisting of minimum and maximum values is constructed. Each interval corresponds to a physical feature term and has a clear numerical boundary. The fusion feature values in the initial judgment label are compared with the maturity confidence interval. When all fusion values fall within the corresponding confidence interval, they are judged to be consistent; otherwise, a re-observation strategy is triggered.
6. The automatic harvesting and grading method for fruits and vegetables based on AI vision and dynamic control according to claim 5, characterized in that, If the stability of the maturity confidence interval does not reach the preset threshold, the re-observation strategy will proceed as follows: Determine whether the fluctuation range of key features in the maturity confidence interval exceeds the set threshold. When the fluctuation of the gray value of the red channel is greater than 8 or the fluctuation of the sugar spot area exceeds 3% of the peel area, perform a supplementary shooting operation, control the image acquisition device to finely adjust the spatial position and pitch angle based on the current posture, and acquire supplementary images to complete the key maturity feature area. By combining the three-dimensional spatial position and posture information of the fruit in the re-shot images, the current accessibility and gripping conditions of the gripper are evaluated, including the distance between the gripping center and the centroid of the fruit, the opening width of the gripper, the gripping angle tolerance, and the occlusion interference in the gripping path, to determine whether the picking action can be performed. The re-shot images are incorporated into the dataset under the original ID number. The maturity features and confidence intervals are updated. If all key features meet the stability requirements of the confidence intervals, the picking action path and execution sequence are regenerated based on the clamping state evaluation results. If the requirements are still not met, automatic picking is stopped and anomalies are marked.
7. The automatic fruit and vegetable harvesting and grading method based on AI vision and dynamic control according to claim 6, characterized in that, After harvesting, the results are fed back into the feature database. The perspective sensitivity index is adjusted based on the difference between the harvested and predicted results, and the feature database is dynamically updated. The closed-loop optimization steps for recognition control are as follows: Collect actual results information during the picking stage, including whether the clamping was successfully completed, whether there was slippage or accidental grasping, whether the fruit was intact, and whether the ripeness was confirmed by manual or sensor to be consistent with the prediction results. The information is then compared with the predicted features according to the unique number of the fruit to identify the image perspective and corresponding feature items with significant residuals. Based on the image number and imaging parameters that have errors, the mapping area of the corresponding viewpoint in the three-dimensional surface model of the fruit is extracted, and the viewpoint sensitivity index value of the mapped area in the multi-angle spectral texture feature library is updated to reflect the poor recognition stability of this viewpoint in the current environment. The corrected sensitivity index is used to reconstruct the visibility prediction map, and the camera pose arrangement and feature fusion weight strategy in subsequent image acquisition are adjusted accordingly to ensure that the recognition results are more stable in subsequent tasks. The picking path and execution parameters are continuously optimized to build a feedback-driven closed-loop update mechanism.