Visual guidance sorting control system and method based on target recognition
By acquiring the view gradient distribution image in the vision-guided sorting control method and performing membership stitching and compensation registration, combined with multi-view information, the problem of inaccurate identification of disordered stacked items in the prior art is solved, and high-precision item classification and sorting is achieved.
Patent Information
- Application Number
- CN202511092529.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-14
AI Technical Summary
Existing visually guided sorting control methods based on target recognition rely on single-view images or two-dimensional features, resulting in inaccurate recognition of the spatial structure and stacking relationship of items in disordered stacking states, which cannot meet the needs of high-precision, multi-target automatic sorting in complex scenarios.
By acquiring the view gradient distribution image of the items to be sorted in a disordered stacked state, membership stitching is performed to generate membership parallax labels. The sorting description index is determined by combining the real-time acquired guidance morphological features, and compensation registration is performed to generate the trajectory of neighboring pixels. Based on the sorting description index and corner fitting sequence, multi-view target classification and recognition are performed to achieve accurate positioning and classification of the sorting guidance area.
It improves the accuracy of spatial perception guidance for disordered stacked items, enhances the recognition accuracy and robustness of sorting areas, supports high-precision spatial geometry analysis and path planning, realizes accurate classification and recognition of multi-category items in complex scenarios, and improves the intelligence level of vision-guided sorting.
Smart Images

Figure CN120953693A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual recognition guidance technology, and more specifically, to a visual guidance sorting control system and method based on target recognition. Background Technology
[0002] Visual recognition guidance refers to the process of integrating technologies such as computer vision, deep learning, and 3D reconstruction to identify, locate, and extract features of items to be sorted in real time, providing precise operational guidance information for automated sorting systems. In this process, the system uses multi-view cameras, depth sensors, and other devices to acquire multimodal visual data of items in the stacking area or on the conveyor line. Through algorithms such as object detection, instance segmentation, and pose estimation, it identifies the category, spatial location, orientation, and boundary contours of the items. Visual recognition guidance can not only identify item categories but also output detailed spatial location information and geometric features to guide robotic arms or mobile gripping devices to accurately plan paths, adjust gripping angles, and control force, achieving precise sorting, handling, or packing operations for different items.
[0003] However, existing visual guidance sorting control methods based on target recognition generally rely on single-view images or two-dimensional features, leading to inaccurate identification of the spatial structure and stacking relationships of items in disordered stacking states. This results in significant errors in the positioning and classification of the sorting guidance area, failing to meet the high-precision, multi-target automated sorting requirements in complex scenarios, thus limiting sorting efficiency and intelligence levels. Therefore, how to achieve comprehensive spatial perception guidance for disordered stacking items through multi-view image fusion and depth information integration to improve the accuracy of visual guidance sorting is a challenge facing the industry. Summary of the Invention
[0004] This application provides a visually guided sorting control system and method based on target recognition, which can provide comprehensive spatial perception guidance for disordered stacked items by combining multi-view image fusion and depth information, thereby improving the accuracy of visually guided sorting.
[0005] In a first aspect, this application provides a visually guided sorting control method based on target recognition, the control method comprising the following steps: Obtain the view gradient distribution image of the items to be sorted in a disordered stacked state; The view gradient distribution image is subjected to membership stitching to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image. Based on the membership disparity labels and the real-time acquired guidance morphology features, the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined. Determine the stacked pose image of the items to be sorted, perform compensation registration on the stacked pose image, generate the neighboring pixel trajectory of dynamic perception vision in the sorting guidance area, and then determine the corner point fitting sequence of the guide mark boundary during classification and sorting based on the neighboring pixel trajectory. Based on the sorting description index and the corner fitting sequence, the categories of dynamically perceived vision in the sorting guidance area are classified and identified using multi-view target recognition to obtain the sorting decision attributes during visually guided sorting. Then, the items to be sorted are classified and sorted using the sorting decision attributes.
[0006] In this embodiment, the view gradient distribution image refers to the spatial feature image of the gradient change pattern of the surface of the items to be sorted under multiple viewpoints.
[0007] In this embodiment, the process of performing membership stitching on the view gradient distribution image to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image specifically includes: Extract the depth features and disparity information from the view gradient distribution image; The depth features and the disparity information are matched and aligned to obtain a multi-view disparity map; The membership disparity label of the view gradient distribution image is determined based on the multi-view disparity map.
[0008] In this embodiment, the membership parallax label refers to a label that identifies different object surface regions and their three-dimensional geometric attributes in a multi-view image.
[0009] In this embodiment, determining the sorting description index corresponding to the sorting guidance area in the visual guidance scene based on the membership parallax label and the real-time acquired guidance morphology features specifically includes: The dynamic sorting attributes in the visually guided scene are determined based on the parallax labels belonging to the category; Real-time acquisition of guidance morphological features in visual guidance scenarios; Based on the aforementioned guiding morphological characteristics, determine the region descriptor corresponding to the sorting guidance area in the guiding scenario; The sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the dynamic sorting attributes and the area descriptor.
[0010] In this embodiment, the sorting description index refers to a set of coded information that uniquely identifies and comprehensively describes each sorting area in the visually guided scene.
[0011] In this embodiment, the compensation registration of the stacked pose image to generate the neighboring pixel trajectory of dynamic perception vision in the sorting guidance area specifically includes: Select a reference frame for the stacked pose image and determine the initial spatial transformation parameters from each frame of the stacked pose image to the reference frame. Subpixel-level compensation registration is performed on all initial spatial transformation parameters to obtain a densely registered image of dynamic perception vision in the sorting guidance area; The neighboring pixel trajectories of dynamic perception in the sorting guidance area are determined based on the dense registration image.
[0012] In this embodiment, the corner point fitting sequence for determining the guide marker boundary during classification and sorting based on the neighboring pixel trajectories specifically includes: Motion feature analysis is performed on the trajectories of the neighboring pixels to obtain dynamic fitting features; Obtain the guiding marking boundaries during classification and sorting; The set of guiding corner points is determined based on the guiding marker boundaries; The corner point fitting sequence for guiding the identification boundaries during classification and sorting is determined based on the dynamic fitting features and the guiding corner point set.
[0013] In this embodiment, based on the sorting description index and the corner fitting sequence, multi-view target classification and recognition are performed on the dynamically perceived visual categories in the sorting guidance area to obtain the sorting decision attributes during visually guided sorting, specifically including: Construct a category multi-view feature tensor based on the sorting description index and the corner fitting sequence; The category confidence distribution in the sorting guidance area is determined by the category multi-view feature tensor; Based on the category confidence distribution, category attribute decisions are made for the visually guided sorting area, and sorting decision attributes are output.
[0014] Secondly, this application provides a visually guided sorting control system based on target recognition, used to execute a visually guided sorting control method based on target recognition, the control system comprising: The image acquisition module is used to acquire the view gradient distribution image of the items to be sorted in a disordered stacked state; The membership stitching module is used to perform membership stitching on the view gradient distribution image, generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image, and determine the sorting description index corresponding to the sorting guidance area in the visual guidance scene based on the membership disparity labels and the real-time acquired guidance morphology features. The compensation registration module is used to determine the stacking pose image of the items to be sorted, perform compensation registration on the stacking pose image, generate the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area, and then determine the corner point fitting sequence of the guide mark boundary during classification and sorting based on the neighboring pixel trajectory. The classification and recognition module is used to perform multi-view target classification and recognition on the categories of dynamically perceived vision in the sorting guidance area based on the sorting description index and the corner fitting sequence, to obtain the sorting decision attributes when visually guiding sorting, and then to classify and sort the items to be sorted through the sorting decision attributes.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: A view gradient distribution image of items to be sorted in a disordered stacked state is acquired; the view gradient distribution image is subjected to membership stitching to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image; the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the membership disparity labels and the real-time acquired guidance morphology features; the stacking pose image of the items to be sorted is determined; the stacking pose image is compensated and registered to generate the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area; then the corner point fitting sequence of the guidance mark boundary during classification and sorting is determined by the neighboring pixel trajectory; the category of the dynamic perception vision in the sorting guidance area is classified and identified by multiple views based on the sorting description index and the corner point fitting sequence to obtain the sorting decision attribute during visual guidance sorting; then the items to be sorted are classified and sorted based on the sorting decision attribute.
[0016] Therefore, this application demonstrates that it can improve the recognition accuracy of the sorting area when the positioning and classification errors in the sorting guidance area are large, making it difficult to achieve high-precision multi-target sorting. Specifically, by acquiring the viewpoint gradient distribution image of the items to be sorted in a disordered stacked state, it is possible to achieve comprehensive spatial information acquisition of complex stacking environments, finely describe the geometric features and edge contours of the items, and improve the completeness and accuracy of the overall visual perception of the items. By performing membership stitching on the viewpoint gradient distribution image and generating membership parallax labels, combined with the real-time acquired guidance morphological features, the sorting description index of the sorting guidance area is determined, achieving efficient fusion of multi-view information and precise regionalization. The positioning system enables high-precision matching and region recognition of the surface edge features of complex items, enhancing the accuracy and robustness of the sorting guidance area. By determining the stacked pose image of the items to be sorted and performing compensation registration, it generates the trajectory of neighboring pixels and further determines the corner point fitting sequence of the guide mark boundary, which can meticulously express the stacking structure and surface features of the items, supporting high-precision spatial geometric analysis and path planning. By performing multi-view target classification and recognition based on the sorting description index and corner point fitting sequence, it obtains the sorting decision attributes during visual guidance sorting, realizing accurate classification and recognition of multi-category items in complex scenes, effectively improving the intelligence level of visual guidance sorting.
[0017] In summary, the technical solution adopted in this application can provide comprehensive spatial perception guidance for disordered stacked items by combining multi-view image fusion and depth information, thereby improving the accuracy of visual guidance sorting. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is an exemplary flowchart of a visually guided sorting control method based on target recognition provided in this application; Figure 2 This is a flowchart illustrating the process of determining the sorting description index provided in this application; Figure 3 This is a flowchart illustrating the process of determining the corner fitting sequence provided in this application; Figure 4 This is a modular structure diagram of a vision-guided sorting control system based on target recognition provided in this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] This application provides a visually guided sorting control system and method based on target recognition. The core of this system is to acquire a viewpoint gradient distribution image of items to be sorted in a disordered stacked state; perform membership stitching on the viewpoint gradient distribution image to generate membership disparity labels corresponding to the upper edge feature points of the viewpoint gradient distribution image; determine the sorting description index corresponding to the sorting guidance area in the visually guided scene based on the membership disparity labels and real-time acquired guidance morphological features; determine the stacking pose image of the items to be sorted; perform compensation registration on the stacking pose image to generate neighboring pixel trajectories of dynamically perceived vision in the sorting guidance area; and then determine the corner point fitting sequence of the guidance mark boundary during classification and sorting based on the neighboring pixel trajectories; perform multi-view target classification and recognition on the categories of dynamically perceived vision in the sorting guidance area based on the sorting description index and the corner point fitting sequence to obtain the sorting decision attributes during visually guided sorting; and then classify and sort the items to be sorted based on the sorting decision attributes.
[0022] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown in the figure, this is an exemplary flowchart of a vision-guided sorting control method based on target recognition according to this embodiment of the present application. The control method includes the following steps: In step S1, an image of the visual gradient distribution of the items to be sorted in a disordered stacking state is obtained.
[0023] In practice, firstly, multiple fixed cameras or mobile cameras installed at the end of the robotic arm are deployed to collect data from multiple perspectives around the sorting area. Each camera completes the calibration of its internal and external parameters during the system initialization phase to ensure that the images captured by different cameras can be used for subsequent spatial relationship deduction. To prevent images from being occluded, the multi-point shooting path planned by the robotic arm is combined to capture images of the stacked area of the items to be sorted from multiple directions (such as the front, side, and top) to obtain multiple two-dimensional images covering the entire scene. All two-dimensional images of the complete scene are used as the perspective gradient distribution images of the items to be sorted in the disordered stacking state.
[0024] It should be noted that, in this application, the view gradient distribution image refers to the spatial feature image of the gradient change law of the surface of the item to be sorted under multiple viewpoints.
[0025] In step S2, the view gradient distribution image is subjected to membership stitching to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image. Based on the membership disparity labels and the real-time acquired guidance morphology features, the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined.
[0026] In this embodiment, the following steps can be used to perform membership stitching on the view gradient distribution image to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image: Extract the depth features and disparity information from the view gradient distribution image; The depth features and the disparity information are matched and aligned to obtain a multi-view disparity map; The membership disparity label of the view gradient distribution image is determined based on the multi-view disparity map.
[0027] In practice, firstly, each viewpoint gradient distribution image is input into a pre-trained depth estimation network model, such as a deep neural network model based on an encoder-decoder structure. The encoder extracts global feature information of the image, including edges, textures, and local curvature. The decoder maps the global feature information to the depth value corresponding to each pixel, outputting a depth map. Simultaneously, it calculates the relative distance from the camera to the object surface pixel by pixel to obtain the depth features corresponding to the image. Then, by performing epipolar constraint calculations between images from different viewpoints, disparity information between each viewpoint is extracted. Disparity information is represented by converting the horizontal or vertical offset of the same object point in different images into disparity values. Next, the depth maps of different viewpoints are normalized to ensure that the depth values of all images are expressed under a uniform scale and reference plane. Normalization can be performed using least squares error matching. Then, by minimizing the depth difference and disparity offset error corresponding to each pixel, corresponding pixel matching pairs are found. If there is local occlusion or incorrect matching points, corrections are made by introducing global smoothing constraints and edge protection mechanisms. After alignment, a unified multi-view disparity map is generated from the depth and disparity information of all viewpoints. Finally, the multi-view disparity map is divided into regions, and the connected components of different objects or local regions of objects are extracted to ensure the geometric continuity within each region. A unique label number is generated for each region as an index for the disparity label. This index not only includes the object region number, but also the region's three-dimensional coordinate range, average depth, and local surface normal direction attribute information. All the indices are used as the disparity labels for the view gradient distribution image.
[0028] It should be noted that, in this application, depth features refer to the set of spatial attributes formed by the distance between each pixel in the image and the camera; disparity information refers to the relative displacement information of the same object point in images from different viewpoints; multi-view disparity map refers to the image of the correspondence and three-dimensional structural integrity of objects under different viewpoints; membership disparity label refers to the label that identifies the surface regions of different objects and their three-dimensional geometric attributes in multi-view images; edge feature points refer to key spatial points that reflect the changes in the shape and contour of the object surface and reflect the location of local geometric abrupt changes.
[0029] Preferably, in this embodiment, the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the belonging parallax label and the real-time acquired guidance morphology features, with reference to... Figure 2 As shown in the figure, this is a flowchart illustrating the process of determining the sorting description index in some embodiments of this application. In this embodiment, determining the sorting description index can be achieved using the following steps: In step S21, the dynamic sorting attributes in the visually guided scene are determined based on the membership parallax label; In step S22, the guiding morphological features in the visual guidance scene are acquired in real time; In step S23, the region descriptor corresponding to the sorting guidance area in the guidance scene is determined based on the guidance morphology features; In step S24, the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the dynamic sorting attribute and the area descriptor.
[0030] In practice, the process begins by reading the generated parallax labels, which contain the 3D coordinates, surface normals, region number, and geometric attributes of each region. These are then combined with the object's surface tilt angle, depth distribution, local curvature, and region accessibility to analyze the dynamic graspability and sortability attributes of each region in the current scene. For example, regions with flat surfaces, no significant tilt, and no occlusion can be marked as high-priority grasping regions; regions with occlusion or high curvature can be marked as low-priority regions. The stacking height information of the current items in the scene (calculated from depth features) and the relative positions between items (determined by 3D coordinates and region boundary relationships) are combined to generate dynamic sorting attributes in the visually guided scene. Next, a high-resolution industrial camera or a robotic arm end-effector camera is deployed to capture real-time images of the current sorting area. During the shooting process, parameters such as exposure time, light source direction, and white balance must be controlled to reduce image deviations caused by changes in ambient light or reflections. The acquired images are then used as guiding morphological features in the visually guided scene. Then, the real-time acquired guidance morphological features are preprocessed, including noise suppression, contrast enhancement, and edge-preserving filtering, to improve image quality. A semantic segmentation network (e.g., Mask R-CNN based on masked region convolutional neural networks) is then used to extract the contours and surface regions of the graspable areas in the scene. For each region, multi-dimensional features such as color histogram, texture gradient, surface normal direction, edge curvature, and region geometric center are extracted and combined to form a vector set of attributes corresponding to the sorting guidance region in the visual guidance scene. This vector set serves as the region descriptor corresponding to the sorting guidance region in the guidance scene. Finally, for each region, the similarity score between the region's dynamic sorting attributes and the region descriptor is calculated (using weighted cosine similarity or Euclidean distance weighted scoring). The highest-scoring match is selected, and a unique sorting description index is generated for each match, thus obtaining the sorting description index corresponding to the sorting guidance region in the visual guidance scene. This sorting description index includes not only basic information such as region number, spatial coordinates, surface normal, and texture pattern, but also action information such as dynamic priority, graspability score, and predicted grasping posture.
[0031] It should be noted that, in this application, dynamic sorting attributes refer to a comprehensive set of attributes describing the graspability, priority, surface state, and spatial accessibility of each area in the current visual scene; guiding morphological features refer to high-precision image data collected in real time to identify the current state of the sorting operation area; area descriptors refer to a comprehensive set of feature vectors used to fully describe the surface texture, geometric characteristics, morphological features, and spatial information of the sorting guidance area; and sorting description index refers to a set of coded information that uniquely identifies and fully describes each sorting area in the visual guidance scene.
[0032] In step S3, the stacking pose image of the items to be sorted is determined, the stacking pose image is compensated and registered, and the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area is generated. Then, the corner point fitting sequence of the guide mark boundary during classification and sorting is determined by the neighboring pixel trajectory.
[0033] In practice, determining the stacking pose image of the items to be sorted can be achieved as follows: First, multiple cameras or end-effector cameras of a robotic arm are used to simultaneously acquire multi-view images of the stacked area from different angles. Depth features and parallax information of each viewpoint are extracted, and a globally consistent 3D point cloud model is generated through feature point matching and spatial geometric alignment. Then, combining point cloud projection and object surface fitting algorithms, the point cloud is converted into a visualized 3D stacking pose image, clarifying the specific position, orientation, stacking level, and local geometric relationships of each item. This 3D stacking pose image is then used as the stacking pose image of the items to be sorted.
[0034] It should be noted that, in this application, the stacked pose image refers to the image data representation of the stacked position and orientation of the items to be sorted in three-dimensional space.
[0035] In this embodiment, the compensation and registration of the stacked pose image to generate the neighboring pixel trajectory of dynamic perception in the sorting guidance area can be achieved by the following steps: Select a reference frame for the stacked pose image and determine the initial spatial transformation parameters from each frame of the stacked pose image to the reference frame. Subpixel-level compensation registration is performed on all initial spatial transformation parameters to obtain a densely registered image of dynamic perception vision in the sorting guidance area; The neighboring pixel trajectories of dynamic perception in the sorting guidance area are determined based on the dense registration image.
[0036] In practice, firstly, the image with the most comprehensive viewpoint, highest coverage, and least occlusion in the stacked pose image sequence is selected as the reference frame. This can be determined by automatically calculated scene information entropy or region integrity index. For each stacked pose image frame, 3D feature points, such as surface keypoints, edge points, or geometrically significant points, are extracted, and similarity matching is performed using feature descriptors to obtain the corresponding set of point pairs. Based on the corresponding point pairs, the rotation matrix and translation vector are calculated using the least squares optimization method to obtain the initial spatial transformation parameters of each frame relative to the reference frame. Then, based on the photometric consistency principle, each frame is projected into the reference frame coordinate system, and the transformation parameters are updated by minimizing the intensity difference of corresponding pixels after projection. Gradient constraints are introduced to ensure edge information alignment. To achieve sub-pixel accuracy, a bicubic interpolation method is used for the grayscale values between pixels to ensure that pixel intensity information can still be accurately calculated even with subtle changes of less than one pixel. After completing the update and optimization, a densely registered image with each frame strictly aligned with the reference frame is output, thus obtaining the densely registered image of dynamic perception vision in the sorting guidance area. Finally, connectivity analysis is performed on the internal pixels in the sorting guidance area. According to the criterion that the Euclidean distance in three-dimensional space is less than a set threshold, adjacent pixels are gradually combined into chain-like trajectories. Subsequently, the local curvature, surface normal change, and continuous gradient information of each trajectory are calculated to verify its geometric continuity. After the trajectory is generated, the three-dimensional coordinate path of each pixel is strictly recorded from the starting point to the ending point, and finally, the neighboring pixel trajectory describing the surface geometric features of the sorting guidance area is formed. The neighboring pixel trajectory of the surface geometric features of the sorting guidance area is used as the neighboring pixel trajectory of dynamic perception vision in the sorting guidance area.
[0037] It should be noted that, in this application, the sorting guidance area refers to a specific target surface used to guide the robotic arm's grasping operation in a vision-guided scenario; the reference frame for stacked pose images refers to a reference image frame used to align the spatial position and orientation of other frame images; the initial spatial transformation parameters refer to a set of parameters describing the initial rotation and translation relationship of non-reference frame images relative to the reference frame in three-dimensional space; the densely registered image refers to a highly consistent spatial image after all pixels are finely aligned with the reference frame; the neighboring pixel trajectory refers to the set of trajectories formed by continuous and mutually adjacent pixels in three-dimensional space after compensated registration; and dynamic perception vision refers to the quantitative characteristics of the continuous visual perception capability that acquires and analyzes the position, orientation, and state changes of items in real time during the sorting process.
[0038] Preferably, in this embodiment, the corner point fitting sequence for guiding the identification boundary during classification and sorting is determined by the neighboring pixel trajectories, with reference to... Figure 3 As shown in the figure, this is a flowchart illustrating the process of determining the corner fitting sequence in some embodiments of this application. In this embodiment, determining the corner fitting sequence can be achieved using the following steps: In step S31, motion feature analysis is performed on the trajectory of the neighboring pixels to obtain dynamic fitting features; In step S32, the guide marking boundaries for sorting and classification are obtained; In step S33, the set of guiding corner points is determined according to the guiding identifier boundary; In step S34, the corner fitting sequence of the guide marker boundary during classification and sorting is determined based on the dynamic fitting features and the guide corner point set.
[0039] In practical implementation, firstly, the changes of pixels contained in each trajectory within the neighboring pixel trajectory in three-dimensional space are analyzed. The motion vector of the trajectory and its changing trend are calculated. Temporal convolutional neural networks or motion estimation methods based on Kalman filtering can be used to extract dynamic feature parameters such as velocity and acceleration of the trajectory, capturing local curvature changes and abrupt changes. Then, combined with curvature calculation and gradient change detection, key turning points and discontinuous motion regions in the trajectory are identified, generating a dynamically fitted feature vector describing the dynamic changes of the trajectory. Next, a depth camera (such as a structured light camera or a time-of-flight camera) is used to acquire the guidance marker boundaries of the current sorting area in real time. A synchronous triggering mechanism ensures that the depth image acquisition is synchronized with the visual guidance scene. Filtering algorithms (such as bilateral filtering) are applied to remove noise and improve the quality of the depth image. After acquisition, the depth image is converted to grayscale, and this grayscale is used as the guidance marker boundaries for classification and sorting. Then, multi-scale corner detection algorithms, such as Harris corner detection combined with the improved Shi-Tomasi algorithm, are used to extract significant corners from the guide marker boundaries. To improve the accuracy of corner detection, depth map gradient information and spatial curvature are integrated. By calculating the gradient change and second derivative of the depth image, corners with significant and stable depth changes are selected, and all selected corners are used as the guide corner set. Finally, the guide corner set is further filtered to exclude abnormal corners that do not match the trajectory dynamic features. Then, the least squares fitting method is used to connect the selected corners sequentially according to the movement trend of the dynamic fitting features to form a fitting sequence. This fitting sequence is smoothed using a curve fitting algorithm (such as cubic spline interpolation) to maintain the spatial continuity and curvature rationality between corners, ultimately forming the corner fitting sequence for the guide marker boundaries during classification and sorting.
[0040] It should be noted that, in this application, the guiding mark boundary refers to the spatial boundary mark that clearly defines the boundary range of the sorting target item to guide the precise grasping and classification operation; the dynamic fitting feature refers to the feature vector that describes the change of motion state and curvature change of the trajectory of neighboring pixels in three-dimensional space; the guiding corner point set refers to the set of significant corner points of the surface detected and located in the guiding mark boundary; and the corner point fitting sequence refers to the information describing the critical path of the sorting target surface.
[0041] In step S4, the categories of dynamically perceived vision in the sorting guidance area are classified and identified using the sorting description index and the corner fitting sequence to obtain the sorting decision attributes during visually guided sorting. Then, the items to be sorted are sorted using the sorting decision attributes.
[0042] In this embodiment, the sorting decision attributes during visually guided sorting are obtained by performing multi-view target classification and recognition on the categories of dynamically perceived vision in the sorting guidance area based on the sorting description index and the corner fitting sequence, using the following steps: Construct a category multi-view feature tensor based on the sorting description index and the corner fitting sequence; The category confidence distribution in the sorting guidance area is determined by the category multi-view feature tensor; Based on the category confidence distribution, category attribute decisions are made for the visually guided sorting area, and sorting decision attributes are output.
[0043] In practical implementation, firstly, the regional geometric features, texture features, and dynamic sorting attribute information contained in the sorting description index are integrated. Combined with the key path and spatial structure information of the object surface expressed in the corner fitting sequence, a tensor data structure is used to uniformly encode these multi-dimensional features, forming a category multi-view feature tensor. This category multi-view feature tensor contains multiple dimensions: spatial location dimension, corner sequence dimension, viewpoint dimension, and feature channel dimension. During encoding, a 3D convolutional neural network or graph neural network can be used to fuse feature maps from different viewpoints. Then, the constructed category multi-view feature tensor is input into a deep classification network, which can be a multi-view convolutional neural network with a spatial attention mechanism. Internally, the network extracts features through convolutional layers, uses a spatial attention module to strengthen attention to key areas, and dynamically adjusts the weights of each viewpoint using a viewpoint attention mechanism. The network ultimately outputs a multi-category confidence distribution vector, where each element represents the probability value of a sample belonging to the corresponding category. This confidence distribution vector serves as the category confidence distribution for dynamic visual perception in the sorting guidance area. Finally, based on the category confidence distribution, a confidence threshold is set to determine the category to which a sample belongs. If the maximum confidence exceeds the threshold, the corresponding category is used as the classification result for the visual guidance area. For cases with low confidence or similar confidence levels across multiple categories, a fusion inference mechanism, such as weighted voting or Bayesian inference, is employed to comprehensively judge the classification by combining historical sorting data and contextual information, thus avoiding misclassification. Ultimately, the determined category labels, confidence values, and related auxiliary information are encoded as sorting decision attributes.
[0044] It should be noted that in this application, the category multi-view feature tensor refers to a multi-dimensional feature set that uniformly encodes the spatial structure features, texture information, and corner point sequences of the target area under multiple perspectives; the category confidence distribution refers to the confidence of the dynamic perception vision in the sorting guidance area belonging to different categories; and the sorting decision attribute refers to the comprehensive recognition result used to describe the target item category and its confidence information.
[0045] In addition, in specific implementation, the sorting decision attributes can be used to classify and sort items as follows: First, the sorting decision attributes are transmitted to the sorting control system. These attributes include item category labels and corresponding confidence information. Based on the category labels, combined with pre-set classification rules and robotic arm gripping strategies, the system automatically matches the corresponding gripping path and motion parameters. The robotic arm control unit adjusts the gripping posture, force, and gripping sequence of the end effector according to the sorting decision attributes to ensure accurate and efficient gripping and handling of different categories of items. At the same time, the system provides real-time feedback on gripping status and item position information, and uses sensors to detect whether the gripping was successful.
[0046] Therefore, this application demonstrates that it can improve the recognition accuracy of the sorting area when the positioning and classification errors in the sorting guidance area are large, making it difficult to achieve high-precision multi-target sorting. Specifically, by acquiring the viewpoint gradient distribution image of the items to be sorted in a disordered stacked state, it is possible to achieve comprehensive spatial information acquisition of complex stacking environments, finely describe the geometric features and edge contours of the items, and improve the completeness and accuracy of the overall visual perception of the items. By performing membership stitching on the viewpoint gradient distribution image and generating membership parallax labels, combined with the real-time acquired guidance morphological features, the sorting description index of the sorting guidance area is determined, achieving efficient fusion of multi-view information and precise regionalization. The positioning system enables high-precision matching and region recognition of the surface edge features of complex items, enhancing the accuracy and robustness of the sorting guidance area. By determining the stacked pose image of the items to be sorted and performing compensation registration, it generates the trajectory of neighboring pixels and further determines the corner point fitting sequence of the guide mark boundary, which can meticulously express the stacking structure and surface features of the items, supporting high-precision spatial geometric analysis and path planning. By performing multi-view target classification and recognition based on the sorting description index and corner point fitting sequence, it obtains the sorting decision attributes during visual guidance sorting, realizing accurate classification and recognition of multi-category items in complex scenes, effectively improving the intelligence level of visual guidance sorting.
[0047] In summary, the technical solution adopted in this application can provide comprehensive spatial perception guidance for disordered stacked items by combining multi-view image fusion and depth information, thereby improving the accuracy of visual guidance sorting.
[0048] Example 2: This application provides a vision-guided sorting control system based on target recognition, referencing... Figure 4 As shown, this figure is a block structure diagram of a vision-guided sorting control system based on target recognition according to this embodiment of the present application. The control system includes: Image acquisition module 100 is used to acquire the view gradient distribution image of the items to be sorted in a disordered stacking state; The membership stitching module 200 is used to perform membership stitching on the view gradient distribution image, generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image, and determine the sorting description index corresponding to the sorting guidance area in the visual guidance scene based on the membership disparity labels and the real-time acquired guidance morphology features. The compensation and registration module 300 is used to determine the stacking pose image of the items to be sorted, perform compensation and registration on the stacking pose image, generate the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area, and then determine the corner point fitting sequence of the guide mark boundary during classification and sorting by the neighboring pixel trajectory. The classification and recognition module 400 is used to perform multi-view target classification and recognition on the categories of dynamically perceived vision in the sorting guidance area based on the sorting description index and the corner fitting sequence, to obtain the sorting decision attributes when visually guiding sorting, and then to classify and sort the items to be sorted through the sorting decision attributes.
[0049] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0050] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0051] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
Claims
1. A visually guided sorting control method based on target recognition, characterized in that, The control method includes the following steps: Obtain the view gradient distribution image of the items to be sorted in a disordered stacked state; The view gradient distribution image is subjected to membership stitching to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image. Based on the membership disparity labels and the real-time acquired guidance morphology features, the sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined. Determine the stacked pose image of the items to be sorted, perform compensation registration on the stacked pose image, generate the neighboring pixel trajectory of dynamic perception vision in the sorting guidance area, and then determine the corner point fitting sequence of the guide mark boundary during classification and sorting based on the neighboring pixel trajectory. Based on the sorting description index and the corner fitting sequence, the categories of dynamically perceived vision in the sorting guidance area are classified and identified using multi-view target recognition to obtain the sorting decision attributes during visually guided sorting. Then, the items to be sorted are classified and sorted using the sorting decision attributes.
2. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The aforementioned view gradient distribution image refers to a spatial feature image of the gradient change pattern on the surface of the items to be sorted from multiple viewpoints.
3. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The process of performing membership stitching on the view gradient distribution image to generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image specifically includes: Extract the depth features and disparity information from the view gradient distribution image; The depth features and the disparity information are matched and aligned to obtain a multi-view disparity map; The membership disparity label of the view gradient distribution image is determined based on the multi-view disparity map.
4. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The aforementioned membership parallax label refers to a label that identifies different object surface regions and their three-dimensional geometric attributes in multi-view images.
5. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the parallax label and the real-time acquired guidance morphology features. Specifically, it includes: The dynamic sorting attributes in the visually guided scene are determined based on the parallax labels belonging to the category; Real-time acquisition of guidance morphological features in visual guidance scenarios; Based on the aforementioned guiding morphological characteristics, determine the region descriptor corresponding to the sorting guidance area in the guidance scenario; The sorting description index corresponding to the sorting guidance area in the visual guidance scene is determined based on the dynamic sorting attributes and the area descriptor.
6. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The sorting description index refers to the set of coded information that uniquely identifies and comprehensively describes each sorting area in the visually guided scenario.
7. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The process of compensating and registering the stacked pose image to generate the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area specifically includes: Select a reference frame for the stacked pose image and determine the initial spatial transformation parameters from each frame of the stacked pose image to the reference frame. Subpixel-level compensation registration is performed on all initial spatial transformation parameters to obtain a densely registered image of dynamic perception vision in the sorting guidance area; The neighboring pixel trajectories of dynamic perception in the sorting guidance area are determined based on the densely registered image.
8. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, The corner fitting sequence for determining the guide marking boundaries during classification and sorting based on the trajectories of neighboring pixels specifically includes: Motion feature analysis is performed on the trajectories of the neighboring pixels to obtain dynamic fitting features; Obtain the guiding marking boundaries during classification and sorting; The set of guiding corner points is determined based on the guiding marker boundaries; The corner point fitting sequence for guiding the identification boundaries during classification and sorting is determined based on the dynamic fitting features and the guiding corner point set.
9. The visually guided sorting control method based on target recognition as described in claim 1, characterized in that, Based on the sorting description index and the corner fitting sequence, multi-view target classification and recognition are performed on the dynamically perceived visual categories in the sorting guidance area to obtain the sorting decision attributes during visually guided sorting, which specifically include: Construct a category multi-view feature tensor based on the sorting description index and the corner fitting sequence; The category confidence distribution in the sorting guidance area is determined by the category multi-view feature tensor; Based on the category confidence distribution, category attribute decisions are made for the visually guided sorting area, and sorting decision attributes are output.
10. A visually guided sorting control system based on target recognition, used to execute a visually guided sorting control method based on target recognition as described in any one of claims 1 to 9, characterized in that, The control system includes: The image acquisition module is used to acquire the view gradient distribution image of the items to be sorted in a disordered stacked state; The membership stitching module is used to perform membership stitching on the view gradient distribution image, generate membership disparity labels corresponding to the upper edge feature points of the view gradient distribution image, and determine the sorting description index corresponding to the sorting guidance area in the visual guidance scene based on the membership disparity labels and the real-time acquired guidance morphology features. The compensation registration module is used to determine the stacking pose image of the items to be sorted, perform compensation registration on the stacking pose image, generate the neighboring pixel trajectory of the dynamic perception vision in the sorting guidance area, and then determine the corner point fitting sequence of the guide mark boundary during classification and sorting based on the neighboring pixel trajectory. The classification and recognition module is used to perform multi-view target classification and recognition on the categories of dynamically perceived vision in the sorting guidance area based on the sorting description index and the corner fitting sequence, to obtain the sorting decision attributes when visually guiding sorting, and then to classify and sort the items to be sorted through the sorting decision attributes.