3D visual guidance positioning method and device for bill unstacking and electronic equipment
Through the 3D visually guided positioning method, the top and side image analysis are used to construct three-dimensional point cloud data of banknote stacks, solving the positioning inaccuracy problem of existing banknote de-palletization systems and achieving efficient and accurate banknote de-palletization.
Patent Information
- Application Number
- CN202510531928.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The existing banknote de-palletization system relies on simple mechanical structures and sensor technology, resulting in inaccurate positioning, incomplete de-palletization or error damage to banknotes, and its accuracy is not high.
Using the 3D visual guide positioning method, the top image of the banknote stack is obtained through the first camera, the position and height information is determined, the cross-sectional profile is extracted using the Blob analysis method, and the side image is obtained by the second camera to construct three-dimensional point cloud data, perform coarse positioning and precise positioning, and obtain the movement orientation and coordinate information of the robot.
It realizes high accuracy and high efficiency of banknote destacking, improves the accuracy and stability of manipulator grabbing, and reduces the risk of banknote damage.
Smart Images

Figure CN120451264A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and specifically to a 3D vision-guided positioning method, device, and electronic equipment for destacking banknotes. Background Art
[0002] With the continuous advancement of industrial automation and intelligentization, the demand for efficient and precise automated equipment is growing across various production and processing fields. Processing large quantities of banknotes requires not only high efficiency but also extremely high accuracy and reliability to ensure the correctness and security of transactions. As a critical step in the banknote processing process, the level of automation in banknote destacking directly impacts the efficiency and quality of the entire banknote processing system.
[0003] Currently, traditional destacking systems rely on simple mechanical structures and basic sensor technology. These systems often operate only under very limited conditions and have strict requirements on the position, shape, and size of the banknote stack. These limitations lead to frequent misalignment of the robotic arm, incomplete destacking, and even accidental damage to banknotes. Consequently, related art banknote destacking methods suffer from low destacking accuracy.
[0004] Therefore, there is an urgent need for a 3D vision-guided positioning method, device and electronic equipment for banknote destacking. Summary of the Invention
[0005] The present application provides a 3D vision-guided positioning method, device, and electronic equipment for banknote destacking, thereby improving the accuracy of banknote destacking.
[0006] In a first aspect of the present application, a 3D vision-guided positioning method for destacking banknotes is provided, the method comprising: acquiring a top image of the banknote stack through a first camera located directly above the banknote stack; determining position information and height information of the banknote stack based on the top image of the banknote stack; determining cross-sectional information of the banknote stack based on the position information and the height information; analyzing the top image using a Blob analysis method based on the cross-sectional information to obtain a cross-sectional profile of the banknote stack; acquiring a side image of the banknote stack through a second camera located on the side of the banknote stack; analyzing the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack; performing coarse positioning on the three-dimensional point cloud data to obtain a moving position of a manipulator of a destacking device, and performing fine positioning based on the moving position to obtain coordinate information of the movement of the manipulator, the coordinate information including X coordinate, Y coordinate, Z coordinate and rotation angle.
[0007] By adopting the above technical solution, the first camera first captures an image of the top of the banknote stack. Based on this image, the stack's position and height are determined, and the cross-sectional information of the stack is obtained. The top image is then analyzed using Blob analysis to extract the stack's cross-sectional profile. A second camera then captures a side image of the stack, which is analyzed based on the cross-sectional profile to generate three-dimensional point cloud data of the stack. Finally, the three-dimensional point cloud data is subjected to coarse and fine positioning to determine the manipulator's movement orientation and precise coordinate information. This method fully utilizes visual information from the top and sides of the banknote stack. Through image analysis and point cloud processing, precise control of the manipulator's grasping position and angle is achieved, significantly improving the accuracy and efficiency of banknote destacking.
[0008] Optionally, determining the cross-sectional information of the stack of banknotes based on the position information and the height information specifically includes: obtaining the highest point in the top image of the stack of banknotes; taking the highest point as a reference, subtracting a preset height value to obtain a target cross-sectional height; obtaining the cross-sectional profile of the stack of banknotes at the target cross-sectional height to obtain initial cross-sectional information; performing image enhancement processing on the initial cross-sectional information to eliminate the proliferative area in the initial cross-sectional information to obtain the cross-sectional information of the stack of banknotes.
[0009] By employing this technical solution, the highest point in the image of the banknote stack's top is first captured. Using this highest point as a reference, a preset height value is subtracted to obtain the target cross-sectional height. The banknote stack's cross-sectional profile is then captured at the target cross-sectional height to obtain initial cross-sectional information. Finally, image enhancement processing is performed on this initial cross-sectional information to eliminate any overgrowth within the initial cross-sectional information, resulting in the final cross-sectional information of the banknote stack. This method fully accounts for the uneven height of the banknote stack's top. By selecting the highest point as a reference point and moving it down a preset height to obtain the target cross-sectional height, the representativeness and accuracy of the cross-sectional information are ensured. Furthermore, image enhancement processing eliminates overgrowth within the cross-sectional profile, further improving the reliability of the cross-sectional information.
[0010] Optionally, based on the cross-sectional information, the top image is analyzed using a Blob analysis method to obtain the cross-sectional profile of the stack of banknotes, specifically including: binarizing the top image to obtain a binary image; marking connected domains on the binary image to obtain multiple connected domains; performing area screening on the multiple connected domains to determine a target connected domain whose area is greater than or equal to a preset area threshold; and performing contour extraction on the target connected domain to obtain the cross-sectional profile of the stack of banknotes.
[0011] By adopting the above technical solution, the top image of the banknote stack is first binarized to obtain a binary image. The binary image is then labeled for connected domains to obtain multiple connected domains. The connected domains are then screened for area to identify target connected domains whose areas are greater than or equal to a preset area threshold. Finally, contour extraction is performed on the target connected domains to obtain the cross-sectional profile of the banknote stack. This method uses binarization to separate the banknote stack area from the background area. Connected domain labeling and area screening accurately locate the effective area of the banknote stack, automatically eliminating background noise and small interfering areas, and significantly improving the accuracy and efficiency of cross-sectional profile extraction. Furthermore, a contour extraction algorithm is used to obtain a vector description of the cross-sectional profile, providing important geometric constraints for subsequent 3D point cloud generation.
[0012] Optionally, the side image is analyzed based on the cross-sectional profile to obtain three-dimensional point cloud data of the stack of banknotes, specifically including: projecting the cross-sectional profile onto the side image to obtain a projection profile; extracting a side profile that matches the projection profile on the side image based on the projection profile; constructing a three-dimensional model of the stack of banknotes based on the side profile and the cross-sectional profile, and extracting three-dimensional point cloud data of the stack of banknotes based on the three-dimensional model.
[0013] By adopting the above technical solution, the cross-sectional profile of the banknote stack is first projected onto the side image to obtain the projected profile. Then, using the projected profile as a reference, a matching side profile is extracted from the side image. A three-dimensional model of the banknote stack is then constructed based on the side and cross-sectional profiles. Finally, three-dimensional point cloud data of the banknote stack is extracted from the three-dimensional model. This method cleverly combines the two-dimensional cross-sectional profile with the side image, obtaining complete three-dimensional profile information of the banknote stack through cross-sectional profile projection and side profile extraction. Based on this, a three-dimensional model of the banknote stack is constructed and three-dimensional point cloud data is extracted from it, fully utilizing the geometric constraints between the cross-sectional and side profiles of the banknote stack. The three-dimensional point cloud data generated by this method accurately depicts the spatial shape and size of the banknote stack, providing a reliable decision-making basis for the robot's motion planning and grasping control.
[0014] Optionally, the coarse positioning of the three-dimensional point cloud data to obtain the moving position of the manipulator of the destacking device specifically includes: slicing the three-dimensional point cloud data to obtain multiple point cloud slices; performing feature extraction on the multiple point cloud slices to obtain feature vectors of each point cloud slice; determining the optimal destacking slice based on the feature vector, and determining the moving position of the manipulator according to the optimal destacking slice.
[0015] By adopting the above technical solution, the three-dimensional point cloud data is first sliced to obtain multiple point cloud slices. Feature extraction is then performed on the point cloud slices to obtain a feature vector for each point cloud slice. Based on the feature vectors, the optimal destacking slice is determined, and the movement direction of the robot arm is determined based on this slice. This method fully utilizes the spatial distribution characteristics of the three-dimensional point cloud data of the banknote stack. Through slicing, the three-dimensional point cloud is converted into a series of two-dimensional slices and the feature vectors of the slices are extracted, achieving a simplified and structured representation of the point cloud data. On this basis, through feature vector analysis and comparison, the optimal destacking slice is determined—that is, the slice most suitable as the starting position for the robot to destacking—and the movement direction of the robot arm is further determined. Slicing and feature extraction greatly simplify the calculation process and improve positioning efficiency. At the same time, by selecting the optimal destacking slice as the starting point for the robot arm to move, the success rate of destacking can be effectively improved and the risk of damage to the banknotes can be reduced.
[0016] Optionally, the precise positioning is performed based on the moving orientation to obtain the coordinate information of the movement of the manipulator, specifically including: acquiring local three-dimensional point cloud data at the moving orientation; performing plane fitting on the local three-dimensional point cloud data to obtain a fitting plane; and determining the target grasping position of the manipulator based on the fitting plane to obtain the coordinate information of the movement of the manipulator.
[0017] By employing this technical solution, local 3D point cloud data is first acquired in the moving orientation. A plane fitting is then performed on this local point cloud data to produce a fitted plane. Finally, based on the fitted plane, the target grasping position of the manipulator is determined, providing coordinate information for the manipulator's movement. This method further refines and optimizes the manipulator's grasping position based on the moving orientation obtained through coarse positioning, achieving high-precision positioning control. By selecting local point cloud data in the moving orientation and performing plane fitting on it, the local geometric features of the banknote surface are derived, avoiding interference and noise from the global point cloud data.
[0018] Optionally, based on the fitting plane, the target grasping position of the manipulator is determined to obtain the coordinate information of the movement of the manipulator, specifically including: extracting inner points on the fitting plane to obtain a plane point cloud; performing convex hull analysis on the plane point cloud to obtain a convex hull contour; using the geometric center of the convex hull contour as the target grasping position of the manipulator; using the three-dimensional coordinates of the target grasping position as the X-coordinate and Y-coordinate of the manipulator, using the normal vector of the fitting plane as the rotation angle, and obtaining the coordinate information according to a preset Z coordinate.
[0019] By adopting the above technical solution, the inner points are first extracted on the fitting plane to obtain a plane point cloud. The plane point cloud is then subjected to convex hull analysis to obtain the convex hull contour. The geometric center of the convex hull contour is then used as the target grasping position of the manipulator. Finally, the three-dimensional coordinates of the target grasping position are used as the X- and Y-coordinates of the manipulator, the normal vector of the fitting plane is used as the rotation angle, and combined with the preset Z coordinate, complete coordinate information is obtained. By extracting the inner points on the fitting plane, this method eliminates abnormal points and outliers, thereby obtaining a more accurate and stable plane point cloud. On this basis, the convex hull contour of the plane point cloud is obtained through convex hull analysis, fully considering the boundary constraints and security of the grasping position. Using the geometric center of the convex hull contour as the grasping position ensures that the grasping point is in the central area of the banknote surface, improving the stability and reliability of the grasping.
[0020] In a second aspect of the present application, a 3D vision-guided positioning device for banknote destacking is provided, which includes: an acquisition module and a processing module, wherein: the acquisition module is used to acquire a top image of the banknote stack through a first camera located directly above the banknote stack; the processing module is used to determine the position information and height information of the banknote stack based on the top image of the banknote stack; the processing module is also used to determine the cross-sectional information of the banknote stack based on the position information and the height information; the processing module is also used to perform a Blob analysis on the top image based on the cross-sectional information. The processing module is further configured to analyze the side image based on the cross-sectional profile to obtain a cross-sectional profile of the stack of banknotes; the acquisition module is further configured to acquire a side image of the stack of banknotes through a second camera located on the side of the stack of banknotes; the processing module is further configured to analyze the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the stack of banknotes; the processing module is further configured to perform coarse positioning on the three-dimensional point cloud data to obtain a moving position of a manipulator of the destacking device, and perform fine positioning based on the moving position to obtain coordinate information of the movement of the manipulator, wherein the coordinate information includes an X coordinate, a Y coordinate, a Z coordinate, and a rotation angle.
[0021] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, any one of the methods described above is executed.
[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. The first camera captures an image of the top of the banknote stack. Based on this image, the stack's position and height are determined, and cross-sectional information of the stack is obtained. Blob analysis is then used to analyze the top image and extract the stack's cross-sectional profile. A second camera then captures a side image of the stack, which is analyzed based on the cross-sectional profile to generate three-dimensional point cloud data of the stack. Finally, the three-dimensional point cloud data is subjected to coarse and fine positioning to determine the robot's movement orientation and precise coordinate information. This method fully utilizes visual information from the top and sides of the banknote stack. Through image analysis and point cloud processing, precise control of the robot's grasping position and angle is achieved, significantly improving the automation and efficiency of banknote destacking. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flow chart of a 3D vision-guided positioning method for banknote destacking disclosed in an embodiment of the present application; Figure 2 This is a module diagram of a 3D vision-guided positioning method for banknote destacking disclosed in an embodiment of the present application; Figure 3 This is a structural diagram of an electronic device disclosed in an embodiment of the present application.
[0025] Description of the accompanying drawings: 201, acquisition module; 202, processing module; 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0027] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0028] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0029] This application provides a 3D vision guided positioning method for banknote destacking, referring to Figure 1 , Figure 1 This is a flow chart of a 3D vision-guided positioning method for banknote destacking provided in an embodiment of the present application. This method is applied to a server that executes a 3D vision-guided positioning program for banknote destacking. The method includes steps S101 to S107, which are as follows: Step S101: Acquire a top image of the banknote stack through a first camera located directly above the banknote stack.
[0030] In step S101, the server captures an image of the top of the banknote stack using a first camera located directly above the stack. Specifically, the server establishes a communication connection with the first camera, which can be either wired or wireless. The server sends a capture command to the first camera via this communication connection, controlling the first camera to capture the top of the banknote stack. After receiving the capture command from the server, the first camera captures the top of the banknote stack according to preset capture parameters, such as aperture size, focal length, and exposure time, to capture an image of the top of the stack. The captured top image is then transmitted to the server via the communication connection.
[0031] Step S102: Determine the position information and height information of the banknote stack based on the top image of the banknote stack.
[0032] In step S102, the server determines the position and height of the banknote stack based on the image of the banknote stack's top. Specifically, the server performs image analysis on the image of the banknote stack's top acquired in step S101, extracting the position and height information of the banknote stack from the image. First, the server performs preprocessing on the image of the banknote stack's top, such as image denoising and image enhancement, to remove noise interference and improve image quality for subsequent image analysis. The server then performs binarization on the preprocessed image, converting it into a black and white binary image. Using a preset threshold, the banknote stack is separated from the background, resulting in a binary image of the banknote stack. Next, the server performs edge detection on the binary image of the banknote stack to extract the edge contour of the banknote stack. Edge detection algorithms such as the Canny algorithm and the Sobel algorithm can be used. Edge detection can determine the position of the banknote stack in the image, namely, the horizontal and vertical coordinates of the banknote stack. Finally, the server analyzes the edge contour of the banknote stack to determine its height. Specifically, the highest point and the lowest point can be found in the edge contour, and the vertical distance between the two points is the height of the banknote stack.
[0033] For example, suppose the server receives an 800×600 pixel image of the top of a stack of banknotes. After preprocessing and binarization, a binary image of the stack is obtained. Edge detection is performed on the binary image using the Canny algorithm to obtain the edge outline of the stack. Analysis of the edge outline reveals the stack's position in the image as follows: (100, 50) for the upper left corner and (700, 550) for the lower right corner. Furthermore, by finding the highest and lowest points in the edge outline, the stack's height is determined to be 500 pixels. Based on the image size of the top of the stack and camera parameters such as focal length and object distance, pixel coordinates and pixel height can be converted to actual physical coordinates and height.
[0034] Step S103: determining cross-sectional information of the banknote stack according to the position information and the height information.
[0035] In step S103, the cross-sectional information of the stack of banknotes is determined based on the position information and the height information, specifically including: obtaining the highest point in the top image of the stack of banknotes; taking the highest point as a reference, subtracting a preset height value to obtain a target cross-sectional height; obtaining the cross-sectional profile of the stack of banknotes at the target cross-sectional height to obtain initial cross-sectional information; performing image enhancement processing on the initial cross-sectional information to eliminate the proliferative area in the initial cross-sectional information to obtain the cross-sectional information of the stack of banknotes.
[0036] Specifically, the server determines the cross-sectional information of the banknote stack based on its position and height information. Specifically, the server locates the highest point in the image of the banknote stack's top and, using this highest point as a reference, subtracts a preset height value to obtain a target cross-sectional height. The server then extracts the banknote stack's cross-sectional profile at the target cross-sectional height to obtain initial cross-sectional information. Finally, the server performs image enhancement processing on this initial cross-sectional information to remove any overgrowth, resulting in the final banknote stack cross-sectional information.
[0037] First, the server searches for the point with the largest pixel value in the image of the top of the banknote stack—the brightest point in the image—as the highest point in the stack. The vertical coordinate of this highest point is the maximum height of the stack. The server then subtracts a pre-set height value (e.g., 5 mm) from the vertical coordinate of the highest point to obtain the vertical coordinate of the target cross-section, i.e., the target cross-section height. The target cross-section height is selected based on the thickness of the banknotes and the neatness of the stack. Generally, a height slightly lower than the highest point of the stack is selected to prevent the cross-section profile from being affected by warping or wrinkling of the top banknotes. This is not a limitation of this application.
[0038] Next, the server scans the image of the top of the banknote stack horizontally at the target cross-sectional height, extracting the pixel values at the target cross-sectional height. This pixel value distribution along a horizontal line represents the cross-sectional profile of the banknote stack. This profile reflects the shape and dimensions of the banknote stack at the target cross-sectional height. However, due to factors such as reflections and shadows on the banknote stack's surface, the extracted cross-sectional profile may contain some noise or protrusions, known as hyperplasia. These protrusions can affect the accuracy of subsequent cross-sectional profile analysis and 3D point cloud data generation. Therefore, the server performs image enhancement on the initial cross-sectional profile. Image processing algorithms such as median filtering and morphological filtering can be used for image enhancement. Median filtering removes isolated noise points from the cross-sectional profile, while morphological filtering smoothes the profile curve and removes protrusions. Image enhancement removes these protrusions, resulting in a more accurate and smooth cross-sectional profile of the banknote stack.
[0039] Step S104: Based on the cross-sectional information, the top image is analyzed using a Blob analysis method to obtain a cross-sectional profile of the banknote stack.
[0040] In step S104, based on the cross-sectional information, the top image is analyzed using a Blob analysis method to obtain a cross-sectional profile of the banknote stack, specifically including: binarizing the top image to obtain a binary image; marking connected domains on the binary image to obtain multiple connected domains; performing area screening on the multiple connected domains to determine a target connected domain whose area is greater than or equal to a preset area threshold; and performing contour extraction on the target connected domain to obtain a cross-sectional profile of the banknote stack.
[0041] Specifically, the server first binarizes the image of the top of the banknote stack. Binarization involves dividing the image's grayscale values into two values, 0 and 1, based on a threshold, resulting in a binary image consisting solely of black and white. By properly setting the binarization threshold, the banknote stack area can be separated from the background area, resulting in pixels with a value of 1 (white) in the stack area and 0 (black) in the background area. Binarization algorithms can be used, such as fixed threshold methods or the Otsu algorithm.
[0042] The server then performs connected domain labeling on the binary image. A connected domain is a set of interconnected pixels with the same pixel value. Connected domain labeling allows the binary image to be segmented into multiple, disjoint connected domains, each corresponding to an independent object or region. Connected domain labeling algorithms include two-pass scanning and seed filling. After labeling, each connected domain is assigned a unique identification number.
[0043] Next, the server screens the marked connected domains by area. Since the binarized and connected domain-marked image may contain some small noise or background areas, these areas are typically small. To extract the true cross-sectional profile of the banknote stack, the connected domains can be screened based on their area. The server sets a preset area threshold, such as 1000 square pixels, and selects connected domains with an area greater than or equal to this threshold as target connected domains. Connected domains with an area smaller than this threshold are considered noise or background and are removed.
[0044] Finally, the server performs contour extraction on the selected target connected domain. Contour extraction involves obtaining the boundary curve of the target connected domain, i.e., the outer contour of the connected domain. This process yields a closed curve representing the cross-sectional profile of the banknote stack.
[0045] For example, suppose a server binarizes an image of the top of a stack of banknotes, generating an 800×600 pixel binary image. Connected domains are labeled on this binary image, yielding 50 connected domains. After area screening, connected domains with an area smaller than 1000 square pixels are removed, resulting in three target connected domains. Contour extraction is performed on each of these three target connected domains, yielding three closed curves. The closed curve with the largest area represents the true cross-sectional contour of the banknote stack, while the remaining two curves likely correspond to reflective or shadowed areas on the banknote stack's surface.
[0046] Step S105: Acquire a side image of the banknote stack through a second camera located on the side of the banknote stack.
[0047] In step S105, the server uses a second camera located to the side of the banknote stack to capture a side image of the banknote stack. Specifically, the server establishes a communication connection with the second camera, controls the second camera to capture the side image of the banknote stack, and transmits the captured side image to the server for subsequent processing.
[0048] Step S106: Analyze the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack.
[0049] In step S106, the side image is analyzed based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack, specifically including: projecting the cross-sectional profile onto the side image to obtain a projection profile; extracting a side profile that matches the projection profile on the side image based on the projection profile; constructing a three-dimensional model of the banknote stack based on the side profile and the cross-sectional profile, and extracting three-dimensional point cloud data of the banknote stack based on the three-dimensional model.
[0050] Specifically, the server first projects the banknote stack's cross-sectional profile obtained in step S104 onto the side image acquired in step S105, generating a projected profile. The projected profile reflects the corresponding position and shape of the cross-sectional profile on the side image. Because the cross-sectional profile and the side image describe the shape of the banknote stack from different perspectives, projecting the cross-sectional profile onto the side image allows the two pieces of information to be linked, providing a reference for subsequent 3D reconstruction.
[0051] The server then extracts a matching side profile from the side image, using the projected profile as a reference. Specifically, the server searches the side image along the projected profile, looking for an edge profile that matches the shape and position of the projected profile. This edge profile represents the height and edge shape of the banknote stack in the side image. Together with the cross-sectional profile, it provides complete 3D information about the banknote stack.
[0052] Next, the server constructs a 3D model of the banknote stack based on the side and cross-sectional profiles. A 3D model is a digital description of the banknote stack's shape, represented using meshes or voxels. The server aligns the side and cross-sectional profiles in 3D space and, through interpolation and fitting algorithms, generates a continuous 3D surface, forming the 3D model of the banknote stack.
[0053] Finally, the server extracts the 3D point cloud data of the banknote stack from the 3D model. A 3D point cloud is a data set consisting of a large number of discrete 3D coordinate points, each representing a sample point on the object's surface. The server uniformly samples the 3D model, generating a series of 3D coordinate points that form the 3D point cloud data of the banknote stack. This 3D point cloud data digitally records the spatial shape and position of the banknote stack, providing a direct reference for the subsequent motion control of the destacking robot.
[0054] For example, suppose the server obtains a cross-sectional profile consisting of 100 2D coordinate points in step S104 and acquires an 800×600 pixel side profile image in step S105. The server projects the cross-sectional profile vertically onto the side profile image, obtaining a projected profile consisting of 100 2D coordinate points. The server then searches the side profile image for the edge profile that best matches the shape and position of the projected profile, extracting a side profile consisting of 120 2D coordinate points.
[0055] Next, the server aligns the side and cross-sectional profiles in 3D space and, using a triangulation algorithm, generates a 3D mesh model composed of 10,000 triangles, accurately representing the 3D shape of the banknote stack. Finally, the server uniformly samples the 3D mesh model to produce a point cloud dataset consisting of 5,000 3D coordinate points. Each coordinate point represents a sample point on the banknote stack's surface, containing rich spatial shape and position information.
[0056] Step S107: performing coarse positioning on the three-dimensional point cloud data to obtain the moving orientation of the manipulator of the depalletizing device, and performing fine positioning based on the moving orientation to obtain coordinate information of the manipulator's movement, wherein the coordinate information includes X coordinate, Y coordinate, Z coordinate and rotation angle.
[0057] In step S107, the three-dimensional point cloud data is coarsely positioned to obtain the movement orientation of the depalletizing device's manipulator. This specifically includes: slicing the three-dimensional point cloud data to obtain multiple point cloud slices; performing feature extraction on the multiple point cloud slices to obtain feature vectors for each point cloud slice; determining the optimal depalletizing slice based on the feature vectors, and determining the movement orientation of the manipulator based on the optimal depalletizing slice. Based on the movement orientation, fine positioning is performed to obtain coordinate information for the manipulator's movement. This specifically includes: obtaining local three-dimensional point cloud data at the movement orientation; performing plane fitting on the local three-dimensional point cloud data to obtain a fitting plane; and determining the manipulator's target grasping position based on the fitting plane to obtain coordinate information for the manipulator's movement.
[0058] Specifically, the server first performs slicing on the 3D point cloud data obtained in step S106. Slicing involves cutting the 3D point cloud data into multiple 2D planes at equal intervals along a certain direction. Each plane contains a portion of the 3D point cloud data, known as a point cloud slice. Typically, slicing is performed perpendicular to the banknote stack surface, with the slice spacing set based on the banknote thickness. Slicing converts the 3D point cloud data into a series of 2D point cloud slices, facilitating subsequent feature extraction and analysis.
[0059] The server then performs feature extraction on multiple point cloud slices. Feature extraction involves extracting quantitative metrics from point cloud slices that reflect their geometric characteristics, such as the slice's area, perimeter, and centroid coordinates. The server calculates a set of feature quantities for each point cloud slice, forming a feature vector for that slice. This feature vector describes the geometric characteristics of the point cloud slice in a compact numerical form, providing a basis for subsequent selection of the optimal destacking slice.
[0060] Next, the server determines the optimal destacking slice based on the feature vectors of the point cloud slices. This optimal destacking slice is the slice that, among all the point cloud slices, is most suitable as the starting position for the robot to destacker. By analyzing and comparing the feature vectors of the point cloud slices, the server selects the slice with the largest area and the most regular shape as the optimal destacking slice. This is because large, regularly shaped slices typically correspond to intact banknotes in the center of the stack, resulting in a higher destacking success rate. Based on the position of the optimal destacking slice, the server determines the robot's movement direction, specifically, the direction from which the robot should begin destacking the banknotes.
[0061] After determining the manipulator's movement orientation, the server further performs precise positioning to obtain the manipulator's specific coordinate information. The server first acquires local 3D point cloud data at the movement orientation, specifically a subset of 3D point cloud data within a certain range near the optimal destacking slice. The server then performs plane fitting on this local 3D point cloud data, generating a fitted plane that reflects the spatial position and orientation of the banknote surface.
[0062] Finally, based on the fitted plane, the server determines the target grasping position of the manipulator—the 3D coordinates where the manipulator's suction cup should reach. The server decomposes the 3D coordinates of the target grasping position into its X, Y, and Z components. Based on the normal vector direction of the fitted plane, the server calculates the rotation angle of the manipulator's suction cup, forming the complete coordinate information for the manipulator's movement.
[0063] For example, suppose the server slices 3D point cloud data, evenly spaced perpendicular to the banknote stack surface, to produce 100 point cloud slices. For each point cloud slice, three features—area, perimeter, and centroid coordinates—are extracted to form a feature vector. By comparing the feature vectors, the server finds that the 60th point cloud slice has the largest area and the most regular shape. It selects it as the optimal slice for destacking and determines that the robot should begin destacking the banknotes from directly above the stack.
[0064] Next, the server obtains the local 3D point cloud data within a 5mm radius above and below the 60th point cloud slice and performs a plane fitting on it, resulting in a fitted plane with a normal vector orientation of (0.1, 0.2, 0.9). Based on the centroid coordinates of the 60th point cloud slice (100, 200, 50) and the suction cup dimensions (e.g., 30mm), the server determines the target grasping position coordinates as (100, 200, 80). This means the robot's X-coordinates are 100mm, Y-coordinates are 200mm, and Z-coordinates are 80mm. Furthermore, based on the normal vector orientation of the fitted plane, the server calculates that the robot's suction cup needs to rotate 15° around the X-axis and 10° around the Y-axis to maintain parallelism with the banknote surface. At this point, the complete coordinate information for the robot's movement is determined to be (100, 200, 80, 15, 10).
[0065] In a possible embodiment, the target grasping position of the manipulator is determined based on the fitting plane, and the coordinate information of the movement of the manipulator is obtained, which specifically includes: extracting inner points on the fitting plane to obtain a plane point cloud; performing convex hull analysis on the plane point cloud to obtain a convex hull contour; using the geometric center of the convex hull contour as the target grasping position of the manipulator; using the three-dimensional coordinates of the target grasping position as the X-coordinate and Y-coordinate of the manipulator, using the normal vector of the fitting plane as the rotation angle, and obtaining the coordinate information according to a preset Z coordinate.
[0066] Specifically, the server first extracts inliers on the fitting plane to generate a planar point cloud. Inliers are 3D point cloud data points located on or very close to the fitting plane. The server sets a preset distance threshold and labels 3D point cloud data points whose distance from the fitting plane is less than the threshold as inliers, while the remaining points are labeled as outliers. The purpose of extracting inliers is to remove noise and outliers, resulting in more accurate and stable planar point cloud data.
[0067] The server then performs convex hull analysis on the planar point cloud. A convex hull is the smallest convex polygon that can completely contain all of the point cloud data points. In the embodiments of the present application, convex hull analysis can be understood as a computational geometry algorithm that can find the convex hull of a set of points in two-dimensional or three-dimensional space. The server projects the planar point cloud onto a two-dimensional plane and then uses a convex hull algorithm, such as the Graham scan method, to calculate the convex hull of the planar point cloud, resulting in a two-dimensional convex polygon, or convex hull contour. The convex hull contour represents the outer boundary of the planar point cloud and reflects the shape characteristics of the banknote surface.
[0068] Next, the server uses the geometric center of the convex hull as the target grasping position for the robotic arm. The geometric center, also known as the centroid or center of mass, is the arithmetic mean of all points on the convex hull. The server calculates the two-dimensional coordinates of the geometric center by averaging the X and Y coordinates of all points on the convex hull. Using the geometric center as the target grasping position ensures that the robotic arm's suction cup grasps the banknote at the center of the surface, improving grasping stability and reliability.
[0069] Finally, the server maps the 2D coordinates (X, Y) of the target grasping position back into 3D space, obtaining the 3D coordinates (X, Y, Z) of the target grasping position, where the Z coordinate is the Z coordinate of the target grasping position on the fitted plane. The server uses the X and Y coordinates of the target grasping position as the X and Y coordinates of the robot's movement, and the normal vector of the fitted plane as the rotation angle of the robot's suction cup. Based on the preset Z coordinate (i.e., the distance between the suction cup and the banknote surface), the server obtains the complete coordinate information of the robot's movement (X, Y, Z, A, B, C), where A, B, and C represent the rotation angles around the X, Y, and Z axes, respectively.
[0070] For example, suppose the server performs plane fitting on local 3D point cloud data, obtaining a fitted plane with a normal vector of (0.1, 0.2, 0.9). The server extracts inliers on the fitted plane, resulting in a planar point cloud containing 500 3D points. Convex hull analysis is performed on the planar point cloud, yielding a convex hull contour consisting of 20 2D points. The geometric center of the convex hull contour is calculated, yielding the 2D coordinates (50, 80), which are used as the target grasping position for the robot.
[0071] The server then maps the 2D coordinates of the target grasping position (50, 80) back into 3D space and finds the corresponding 3D coordinates (50, 80, 30) on the fitted plane. The server uses the X-coordinates 50 and Y-coordinates 80 as the X and Y coordinates for the robot's movement. It converts the normal vector (0.1, 0.2, 0.9) of the fitted plane into rotation angles (5.7°, 11.5°, 64.2°), sets the distance between the suction cup and the banknote surface to 10 mm, and finally obtains the coordinates for the robot's movement: (50, 80, 40, 5.7, 11.5, 64.2).
[0072] Reference Figure 2 The present application also provides a 3D vision-guided positioning device for banknote destacking, which is a server. The server includes an acquisition module and a processing module, wherein: the acquisition module is used to obtain the top image of the banknote stack through a first camera located directly above the banknote stack; the processing module is used to determine the position information and height information of the banknote stack based on the top image of the banknote stack; the processing module is also used to determine the cross-sectional information of the banknote stack based on the position information and the height information; the processing module is also used to use the Blob analysis method to analyze the top image based on the cross-sectional information. The image is analyzed to obtain a cross-sectional profile of the stack of banknotes; the acquisition module is further used to obtain a side image of the stack of banknotes through a second camera located on the side of the stack of banknotes; the processing module is further used to analyze the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the stack of banknotes; the processing module is further used to perform coarse positioning on the three-dimensional point cloud data to obtain the moving direction of the manipulator of the destacking device, and perform fine positioning based on the moving direction to obtain coordinate information of the movement of the manipulator, wherein the coordinate information includes X coordinate, Y coordinate, Z coordinate and rotation angle.
[0073] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0074] This application also provides an electronic device. Figure 3 , Figure 33. This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0075] The communication bus 302 is used to implement the connection and communication between these components.
[0076] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0077] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0078] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.
[0079] Among them, the memory 305 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also optionally be at least one storage device located away from the aforementioned processor 301. Refer to Figure 3 The memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for a 3D vision-guided positioning method for banknote destacking.
[0080] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 305 for a 3D vision-guided positioning method for destacking banknotes. When executed by one or more processors 301, the electronic device 300 executes one or more of the methods described in the above embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0081] The present application further provides a computer-readable storage medium storing instructions, which, when executed by one or more processors 301 , enable the electronic device 300 to perform one or more of the methods described in the above embodiments.
[0082] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0083] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0084] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0085] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0087] The foregoing is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.
[0088] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. A 3D vision-guided positioning method for banknote destacking, characterized in that: The method comprises: acquiring a top image of the banknote stack by a first camera located directly above the banknote stack; determining position information and height information of the banknote stack based on the top image of the banknote stack; determining cross-sectional information of the banknote stack according to the position information and the height information; Based on the cross-sectional information, the top image is analyzed using a Blob analysis method to obtain a cross-sectional profile of the banknote stack; acquiring a side image of the banknote stack by a second camera located on a side of the banknote stack; Analyzing the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack; The three-dimensional point cloud data is roughly positioned to obtain the moving position of the manipulator of the depalletizing device, and fine positioning is performed based on the moving position to obtain the coordinate information of the manipulator's movement, wherein the coordinate information includes X coordinate, Y coordinate, Z coordinate and rotation angle.
2. The method according to claim 1, characterized in that Determining the cross-sectional information of the banknote stack according to the position information and the height information specifically includes: obtaining the highest point in the top image of the stack of banknotes; Taking the highest point as a reference, subtract the preset height value to obtain the target section height; At the target cross-sectional height, obtaining a cross-sectional profile of the banknote stack to obtain initial cross-sectional information; The initial cross-sectional information is subjected to image enhancement processing to eliminate the proliferation area in the initial cross-sectional information, thereby obtaining the cross-sectional information of the banknote stack.
3. The method according to claim 1, characterized in that The analyzing the top image based on the cross-sectional information using a Blob analysis method to obtain a cross-sectional profile of the banknote stack specifically includes: performing binarization processing on the top image to obtain a binarized image; Marking connected domains on the binary image to obtain multiple connected domains; Performing area screening on the multiple connected domains to determine a target connected domain whose area is greater than or equal to a preset area threshold; Contour extraction is performed on the target connected domain to obtain a cross-sectional contour of the banknote stack.
4. The method according to claim 1, wherein The analyzing the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack specifically includes: Projecting the cross-sectional profile onto the side image to obtain a projected profile; Taking the projection profile as a reference, extracting a side profile that matches the projection profile on the side image; A three-dimensional model of the banknote stack is constructed based on the side profile and the cross-sectional profile, and three-dimensional point cloud data of the banknote stack is extracted based on the three-dimensional model.
5. The method according to claim 1, wherein The coarse positioning of the three-dimensional point cloud data to obtain the moving position of the manipulator of the depalletizing device specifically includes: Slicing the three-dimensional point cloud data to obtain a plurality of point cloud slices; Performing feature extraction on the plurality of point cloud slices to obtain a feature vector of each point cloud slice; Based on the feature vector, an optimal destacking slice is determined, and a moving orientation of the robot arm is determined according to the optimal destacking slice.
6. The method according to claim 1, characterized in that The precise positioning is performed based on the movement orientation to obtain the coordinate information of the movement of the manipulator, specifically including: Acquiring local three-dimensional point cloud data at the moving position; Performing plane fitting on the local three-dimensional point cloud data to obtain a fitting plane; Based on the fitting plane, the target grasping position of the manipulator is determined, and the coordinate information of the movement of the manipulator is obtained.
7. The method according to claim 6, characterized in that Determining the target grasping position of the manipulator based on the fitting plane and obtaining coordinate information of the movement of the manipulator specifically includes: Extracting inliers on the fitting plane to obtain a plane point cloud; Performing convex hull analysis on the plane point cloud to obtain a convex hull contour; Using the geometric center of the convex hull contour as the target grasping position of the manipulator; The three-dimensional coordinates of the target grasping position are used as the X coordinate and Y coordinate of the manipulator, the normal vector of the fitting plane is used as the rotation angle, and the coordinate information is obtained according to the preset Z coordinate.
8. A 3D vision-guided positioning device for banknote destacking, characterized in that: The device comprises an acquisition module (201) and a processing module (202), wherein: The acquisition module (201) is used to acquire a top image of the banknote stack through a first camera located directly above the banknote stack; The processing module (202) is used to determine the position information and height information of the banknote stack based on the top image of the banknote stack; The processing module (202) is further configured to determine cross-sectional information of the banknote stack based on the position information and the height information; The processing module (202) is further configured to analyze the top image using a Blob analysis method based on the cross-sectional information to obtain a cross-sectional profile of the banknote stack; The acquisition module (201) is further configured to acquire a side image of the banknote stack via a second camera located on a side of the banknote stack; The processing module (202) is further configured to analyze the side image based on the cross-sectional profile to obtain three-dimensional point cloud data of the banknote stack; The processing module (202) is further used to perform coarse positioning on the three-dimensional point cloud data, obtain the moving orientation of the manipulator of the depalletizing device, and perform fine positioning based on the moving orientation to obtain coordinate information of the movement of the manipulator, wherein the coordinate information includes X coordinate, Y coordinate, Z coordinate and rotation angle.
9. An electronic device, characterized in that: The electronic device (300) comprises a processor (301), a memory (305), a user interface (303) and a network interface (304), wherein the memory (305) is used to store instructions, the user interface (303) and the network interface (304) are used to communicate with other devices, and the processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device (300) executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.