A real-time three-dimensional reconstruction method and system for the surface of drugs based on binocular vision

By pre-processing and stereo matching of binocular camera images, combined with TSDF algorithm and feature extraction, the real-time and accuracy problems of binocular vision three-dimensional reconstruction are solved, and high-precision reconstruction and accurate capture of drug surfaces are achieved.

CN116758211BActive Publication Date: 2025-07-18TAIYUAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310378285.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-10
Publication Date
2025-07-18
Estimated Expiration
2043-04-10

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology based on binocular vision has poor real-time and low accuracy, resulting in inaccurate drug capture.

Method used

By preprocessing and stereo matching the images collected by the binocular camera, the three-dimensional coordinates and features of the drug are extracted using edge detection and TSDF algorithms, and the TSDF model is updated in combination with global topological invariance and local detail features to obtain the drug surface image.

Benefits of technology

Real-time high-precision reconstruction of the drug surface is achieved, and the accuracy of drug grabbing and position estimation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758211B_ABST
    Figure CN116758211B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for real-time three-dimensional reconstruction of the surface of drugs based on binocular vision, which includes preprocessing the original images collected by a binocular camera. After obtaining the left-eye image and the right-eye image, binocular vision stereo matching is performed to obtain a disparity image; the disparity image is detected to obtain a foreground image containing the target drug; according to the disparity map of the foreground image, the three-dimensional coordinates of the target drug are obtained; according to the three-dimensional coordinates, the disparity map of the foreground image is back-projected into the three-dimensional space by using the TSDF model to obtain a three-dimensional feature volume, and the voxels with TSDF values not greater than 1 in the three-dimensional feature volume are used as target voxels; according to the disparity image, the local detail features and global topological invariance features of the target voxels are extracted; according to the local detail features and global topological invariance features on the foreground image corresponding to the target voxels, the truncation distance and weight of the TSDF model are updated to obtain a TSDF global fusion model; the isosurface with the weighted sum of the directed distances being 0 is calculated to obtain the surface image of the target drug.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and in particular to a method and system for real-time three-dimensional reconstruction of the surface of drugs based on binocular vision. Background Art

[0002] Three-dimensional reconstruction refers to the mathematical process of using two-dimensional projections or images to restore the three-dimensional information (such as shape) of an object, establishing a mathematical model suitable for computer representation and processing of a three-dimensional object, which is the basis for processing, operating, and analyzing its properties in a computer environment, and is also the key technology for establishing a virtual reality that expresses the objective world in a computer. Binocular vision can best simulate the human visual system. Binocular cameras are used to collect images of target objects. According to computer operation analysis, understanding, and recognition of information such as the category and position of the target object, as an auxiliary means, it can be used for target detection and tracking, and can also be used for three-dimensional measurement, etc. Especially in an automated pharmacy, three-dimensional reconstruction is used for robot positioning and grasping in the automated pharmacy. When the robot grasps drugs, a real-time three-dimensional reconstruction model of the drugs is established, which is the basis for positioning and detecting the target drugs. Therefore, whether the drugs can be effectively identified, perceived, and three-dimensionally reconstructed is a prerequisite for the autonomous operation of the automated pharmacy.

[0003] Currently, the real-time performance of three-dimensional reconstruction based on binocular vision is relatively low, and real-time three-dimensional reconstruction based on images relies on RGB-D sensors and is realized by using monocular multi-view images. Without adding sensors, it can be directly used in existing intelligent devices. However, since the RGB-D sensor uses a monocular camera to collect multiple sequence images, there will be a serious problem of repeated inter-frame information. The same pixel point is calculated multiple times, resulting in a large amount of calculation and a high time complexity. Although there is a large overlap between adjacent frames, when calculating the depth map of each frame, it still has to start over, which will also cause the depth maps of adjacent frames calculated to be inconsistent, and the reconstructed results will be very scattered, and even stratification will occur. At the same time, due to reasons such as noise and depth map calculation errors, there will also be some noise points during model reconstruction. Whether it is the phenomenon of model stratification or noise points will affect the reconstructed surface of the drugs, and the accuracy of the reconstructed results will lead to inaccurate detection of the drug grasping points and pose estimation, directly affecting the accuracy of the drug grasping by the drug-taking robot. Therefore, it is necessary to perform high-precision modeling on the drugs. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the problems of poor real-time performance and low accuracy in binocular vision three-dimensional modeling in the prior art.

[0005] To solve the above technical problems, the present invention provides a method for real-time three-dimensional reconstruction of the surface of drugs based on binocular vision, including:

[0006] Preprocess the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image, and perform binocular vision stereo matching to obtain a disparity image;

[0007] Use an edge detection algorithm to detect the disparity image, distinguish the foreground image and the background image, obtain the foreground image containing the target drug, and calculate the three-dimensional coordinates of the corresponding points of the target drug;

[0008] According to the three-dimensional coordinates of the corresponding points of the target drug, use the TSDF algorithm to back-project the disparity map of the foreground image into three-dimensional space to obtain a three-dimensional feature volume, and use the voxels with TSDF values not greater than 1 in the three-dimensional feature volume as target voxels;

[0009] According to the disparity image, extract the contour features, spatial structure features and geometric shape features of the target voxels, and concatenate them to form the original global topological invariance features;

[0010] According to the disparity image, extract the color features, texture features and voxel feature points of the target voxels, and concatenate them to form the original local detail features;

[0011] According to the disparity map of the foreground image, extract the foreground local detail features and foreground global topological invariance features on the foreground image corresponding to the target voxels;

[0012] According to the original global topological invariance features, the original local detail features, the foreground local detail features and the foreground global topological invariance features of each target voxel, update the truncation distance and weight of the TSDF model to obtain the TSDF global fusion model;

[0013] Calculate and obtain the isosurface composed of all target voxels with a weighted sum of directed distances of 0 in the TSDF global fusion model to obtain the surface image of the target drug.

[0014] In an embodiment of the present invention, the preprocessing includes:

[0015] Perform intensity transformation on the pixel points in the left-eye original image and the right-eye original image to obtain a new pixel point intensity value with variables k, a, b, and c to be optimized; take each parameter as a particle in the particle swarm optimization algorithm within a preset range according to a preset step size, and use the particle swarm optimization algorithm update formula to update the velocity and position of the particle until the value of the fitness function converges to obtain the left-eye image and the right-eye image;

[0016] The expression of the intensity transformation is:

[0017]

[0018] Among them, (x, y) represents the coordinates of a pixel point in the original image, f(x, y) represents the image intensity value of this pixel point before preprocessing, m(x, y) represents the local mean, M and N respectively represent the number of rows and columns of the pixel points in the original image, n represents the window size, G m represents the global mean, σ(x, y) represents the local standard deviation, and the parameters k, a, b, and c to be optimized are all unknown parameters. Taking the local mean as the processing center, the original image is preprocessed;

[0019] The particle swarm update formula is as follows:

[0020] V i+1,j (t) = V i,j (t) + c1r 1,j (p i,j - x i,j (t)) + c2r 2,j (p g,j - x i,j (t)),

[0021] x i+1,j (t) = x i,j (t) + v i+1,j (t),

[0022] Among them, t represents the current iteration number, i represents the current particle, j represents the current dimension, V(t) represents the velocity of the particle at the t-th iteration, x(t) represents the position of the particle at the t-th iteration, p i,j represents the historical best solution of the current particle, p g,j represents the best solution of the global particle, c1 is the acceleration coefficient for adjusting p i,j and c2 is the acceleration coefficient for adjusting p g,j , r 1,j ~U(0, 1) and r 2,j ~U(0, 1) represent independent random functions.

[0023] In an embodiment of the present invention, the binocular vision stereo matching is performed to obtain a disparity image, including:

[0024] Using the SURF algorithm to detect feature points in the left-eye image and the right-eye image, and performing sparse representation on the detected feature points;

[0025] Calculating the gradient similarity between the sparse representation feature points of the left-eye image and the sparse representation feature points of the right-eye image as the initial cost of matching;

[0026] Taking the global energy optimal strategy as the cost aggregation function, using WTA to calculate the minimum cost aggregation to obtain the disparity value, and obtaining the disparity image.

[0027] In an embodiment of the present invention, the method for detecting the disparity image by using an edge detection algorithm, distinguishing the foreground image from the background image, and obtaining the foreground image containing the target drug includes:

[0028] Establish a consumption equation for different edges and obtain the optimal solution of the consumption equation;

[0029] Distinguish the foreground image from the background image according to the optimal solution, and obtain the foreground image containing the target drug.

[0030] In an embodiment of the present invention, the process of obtaining the three-dimensional coordinates of the corresponding points of the target drug includes:

[0031] Take a point P in space, f is the focal length of the binocular camera, and b is the baseline, representing the distance between the optical centers of the binocular cameras; P L is the image point of point P under the left-eye camera, and P R is the image point of point P under the right-eye camera, x L is the distance from P L to the imaging plane of the left-eye camera, and x R is the distance from P R to the imaging plane of the right-eye camera. d represents the position difference of the corresponding points of point P in the images formed by the binocular cameras. Then the disparity d is expressed as: d = |x L -x R |;

[0032] The distance between the image point P L under the left-eye camera and the image point P R under the right-eye camera is expressed as:

[0033]

[0034] According to the similarity triangle theory, it can be obtained that:

[0035]

[0036] The disparity d is inversely proportional to Z. If point P is not fixed, the positions of the image points on the left and right images will change as point P moves, and then the disparity will also change. Combining the depth information of the third dimension, the three-dimensional coordinates of point P are expressed as:

[0037]

[0038] In an embodiment of the present invention, the acquisition of the original global topological invariance features includes:

[0039] Taking the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, and extracting the global topological invariance features of the target voxels in the region;

[0040] Obtain the edge contour feature g of the target drug within the region using the snake contour feature extraction algorithm C (v);

[0041] Extract features using the Gabor transform to obtain the first feature vector g gabor (v), and obtain the second feature vector g using the GLCM gray-level histogram feature GLCM (v), the spatial structure feature g of the target drug within the region G (v), denoted as g G (v)=[g gabor (v), g GLCM (v)];

[0042] Detect using the Canny operator to obtain the geometric shape feature g of the target drug within the region P (v);

[0043] Concatenate the edge contour feature, the spatial structure feature, and the geometric shape feature to obtain the global topological invariance feature:

[0044] G(v)=[g C (v); g G (v); g P (v)].

[0045] In an embodiment of the present invention, the acquisition of the local detail features includes:

[0046] Use the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, and take the voxel as the unit with a step size of 1 to extract the local detail features of the target voxels in the region;

[0047] Calculate the color feature T within each target voxel through the mean calculated by the first moment and the variance calculated by the second moment C (v);

[0048] Use the LBP algorithm to extract the texture feature T within each target voxel for distinguishing different drugs LBP (v);

[0049] Use the SURF algorithm to extract the feature points T within each target voxel SURF (v);

[0050] Concatenate the color feature, the texture feature, and the feature points to obtain the local detail features within each target voxel:

[0051] T(v)=[T C (v); T LBP (v); T SURF (v)].

[0052] In one embodiment of the present invention, the acquisition of the TSDF global fusion model includes:

[0053] According to the disparity image of the left-eye image and the right-eye image, extract the original local detail feature T(v) and the original global topological invariance feature G(v) of the target voxel v;

[0054] According to the disparity map of the foreground image, extract the foreground local detail feature T(u) and the foreground global topological invariance feature G(u) on the foreground image corresponding to the target voxel u;

[0055] According to T(v), G(v), T(u) and G(u) of each target voxel, update the truncation distance and weight of the TSDF model;

[0056] The truncation distance update formula is:

[0057] diff(v) = |T(v) - T(u)| + |G(v) - G(u)|;

[0058] The weight update formula is:

[0059] w i (v) = min(w i-1 (v) + λ i (v) + ξ(v), w μ );

[0060] where ξ(v) is the weight update value corresponding to the global topological invariance feature and the local detail feature of each voxel; λ i (v) is the weight increment provided by the target voxel;

[0061] According to the updated truncation distance and weight, obtain the updated TSDF value, expressed as:

[0062]

[0063] According to the updated TSDF value and weight, obtain the TSDF global fusion model

[0064] In one embodiment of the present invention, use the marching cubes algorithm in the TSDF global fusion model to calculate and obtain the isosurface composed of all target voxels with a directed distance weighted sum of 0.

[0065] The embodiment of the present invention also provides a real-time three-dimensional reconstruction system for the surface of a drug based on binocular vision, including:

[0066] A preprocessing module for preprocessing the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image;

[0067] A binocular vision stereo matching module, which is used to perform binocular vision stereo matching on the left-eye image and the right-eye image to obtain a disparity image;

[0068] A target drug segmentation module, which is used to detect the disparity image by using an edge detection algorithm, distinguish the foreground image from the background image, and obtain the foreground image containing the target drug;

[0069] A target drug three-dimensional coordinate acquisition module, which is used to obtain the three-dimensional coordinates of the corresponding points of the target drug according to the disparity map of the foreground image;

[0070] A TSDF model construction module, which is used to inversely project the disparity map of the foreground image into the three-dimensional space by using the TSDF model according to the three-dimensional coordinates to obtain a three-dimensional feature volume, and use the voxels with TSDF values not greater than 1 in the three-dimensional feature volume as target voxels;

[0071] A feature extraction module, which is used to extract the original global topological invariance features and original local detail features of the target voxels according to the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image; according to the disparity map of the foreground image, extract the foreground global topological invariance features and foreground local detail features of the target voxels;

[0072] A TSDF global fusion model reconstruction module, which is used to update the truncation distance and weight according to the relevant features of each target voxel extracted by the feature extraction module to obtain a TSDF global fusion model;

[0073] A drug surface acquisition module, which is used to calculate and obtain the isosurface composed of all target voxels with a weighted sum of directed distances of 0 in the TSDF global fusion model to obtain the target drug surface image.

[0074] The above technical solutions of the present invention have the following advantages compared with the prior art:

[0075] The real-time three-dimensional reconstruction method of the drug surface based on binocular vision according to the present invention uses the disparity image generated after binocular vision matching of the left-eye image and the right-eye image collected by the binocular camera as the region, and extracts the contour features, spatial structure features, and geometric shape features for distinguishing whether the drug is boxed, bottled, or bagged within the region, and concatenates them to form the global topological invariance features; extracts the color features, texture features, and pixel feature points for obtaining information such as the drug type, name, and manufacturer within the region, and concatenates them to form the local detail features; uses the global topological invariance features and local detail features of the corresponding voxels extracted from the target drug in the disparity map of the disparity image and the foreground image to update the truncation distance and weight in the TSDF model, and obtains the TSDF global fusion model; in addition to the truncation distance value and weight for each voxel of the TSDF global fusion model, the global topological invariance features and local detail features of the voxel grid are also added. When calculating the isosurface with a weighted sum of directed distances of 0 using the TSDF global fusion model to obtain the surface of the target drug, a more real-time and accurate three-dimensional reconstruction of the drug surface is realized, providing a more accurate drug recognition model for the drug-taking robot and improving the pose accuracy of grasping the drug. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the drawings, where

[0077] Figure 1 is a flowchart of the steps of the real-time three-dimensional reconstruction method of the drug surface based on binocular vision provided by the present invention;

[0078] Figure 2 is a schematic diagram of the effect of the real-time three-dimensional reconstruction process of the boxed drug surface provided by the present invention;

[0079] Figure 3 is a schematic diagram of the effect of the real-time three-dimensional reconstruction process of the bottled drug surface provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] The following further describes the present invention in conjunction with the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the examples given are not intended to limit the present invention.

[0081] Refer to Figure 1 As shown, the real-time three-dimensional reconstruction method of the drug surface based on binocular vision of the present invention includes:

[0082] S1: Preprocess the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image;

[0083] The preprocessing includes: performing intensity transformation on pixel points in the original image to obtain new pixel point intensity values with the to-be-optimized parameters k, a, b, and c as variables; taking each parameter to take values within a preset range according to a preset step size, and each value is used as a particle in the particle swarm algorithm. Using the particle swarm algorithm update formula, update the velocity and position of the particle until the value of the fitness function converges to obtain a depth image;

[0084] The expression of the intensity transformation is:

[0085]

[0086] where (x, y) represents the coordinates of a pixel point in the original image, f(x, y) represents the image intensity value of this pixel point before preprocessing, m(x, y) represents the local mean, M and N respectively represent the number of rows and columns of the pixel points in the original image, n represents the window size, G m represents the global mean, σ(x, y) represents the local standard deviation, and the to-be-optimized parameters k, a, b, and c are all unknown parameters. Taking the local mean as the processing center, preprocess the original image;

[0087] The particle swarm update formula is:

[0088] V i+1,j (t) = V i,j (t) + c1r 1,j (p i,j -x i,j (t)) + c2r 2,j (p g,j -x i,j (t)),

[0089] x i+1,j (t) = x i,j (t) + v i+1,j (t),

[0090] where t represents the current iteration number, i represents the current particle, j represents the current dimension, V(t) represents the velocity of the particle at the t-th iteration, x(t) represents the position of the particle at the t-th iteration, p i,j represents the historical best solution of the current particle, p g,j represents the best solution of the global particle, c1 is the acceleration coefficient for adjusting p i,j , c2 is the acceleration coefficient for adjusting p g,j , r 1,j ~U(0, 1) and r 2,j ~U(0, 1) represent independent random functions.

[0091] S2: Perform binocular vision stereo matching on the left-eye image and the right-eye image to obtain a disparity image;

[0092] S21: Detect feature points of the left-eye image and the right-eye image using the SURF algorithm, and perform sparse representation on the detected feature points;

[0093] S22: Calculate the gradient similarity between the feature points with sparse representation in the left-eye image and those in the right-eye image as the initial cost of matching;

[0094] S23: Take the global energy optimal strategy as the cost aggregation function, and use WTA to calculate the minimum cost aggregation to obtain the disparity value and acquire the disparity image.

[0095] S3: Detect the disparity image using an edge detection algorithm, distinguish the foreground image from the background image, and acquire the foreground image containing the target drug;

[0096] S31: Establish a cost equation for different edges and find the optimal solution of the cost equation;

[0097] S32: Distinguish the foreground image from the background image according to the optimal solution, and acquire the foreground image containing the target drug.

[0098] S4: Obtain the three-dimensional coordinates of the corresponding points of the target drug according to the disparity map of the foreground image;

[0099] Take a point P in space, f is the focal length of the binocular camera, b is the baseline representing the distance between the optical centers of the binocular cameras; P L is the image point of point P under the left-eye camera, P R is the image point of point P under the right-eye camera, x L is the distance from P L to the imaging plane of the left-eye camera, x R is the distance from P R to the imaging plane of the right-eye camera, d represents the position difference of the corresponding points of point P in the images formed by the binocular cameras, then the disparity d is expressed as: d = |x L - x R |;

[0100] The image point P L under the left-eye camera and the image point P R under the right-eye camera are represented as:

[0101]

[0102] According to the similar triangle theory, it can be obtained that:

[0103]

[0104] The parallax d is inversely proportional to Z. If the point P is not fixed, the positions of the image points on the left and right images will change as P moves during imaging, and thus the parallax will also change. Combining with the depth information of the third dimension, the three-dimensional coordinates of point P are expressed as:

[0105]

[0106] S5: According to the three-dimensional coordinates of the corresponding points of the target drug, use the TSDF algorithm to back-project the parallax map of the foreground image into the three-dimensional space to obtain a three-dimensional feature volume, and use the voxels with TSDF values not greater than 1 in the three-dimensional feature volume as target voxels;

[0107] The binocular camera captures images from multiple perspectives of the target drug, and obtains multiple parallax maps through binocular calculation. Each parallax map is the depth feature of the current image, and the extracted parallax maps are back-projected along each ray into the three-dimensional feature volume representing the image features; the three-dimensional feature volume should completely enclose the reconstructed target drug, and the TSDF value is calculated in the three-dimensional feature volume with voxels as units, and the voxels with TSDF values not greater than 1 are used as target voxels.

[0108] S6: According to the parallax image, extract the contour features, spatial structure features and geometric shape features of the target voxels, and concatenate them to form the original global topological invariance features

[0109] S61: Use the parallax image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, and extract the topological invariance features of the target voxels in the region;

[0110] S62: Use the snake contour feature extraction algorithm to obtain the edge contour feature g C (v);

[0111] S63: Use the gabor transform to extract features to obtain the first feature vector g gabor (v), use the GLCM gray-level histogram feature to obtain the second feature vector g GLCM (v), the spatial structure feature g G (v) of the target drug in the region, which is expressed as g G (v)=[g gabor (v),g GLCM (v)];

[0112] S64: Use the Canny operator to detect and obtain the geometric shape feature g P (v);

[0113] S65: Concatenate the edge contour feature, the spatial structure feature and the geometric shape feature to obtain the topological invariance feature:

[0114] G(v)=[g C (v); g G (v); g P (v)];

[0115] S7: Extract the color feature, texture feature, and voxel feature points of the target voxel according to the disparity image, and concatenate them to form the original local detail feature;

[0116] S71: Take the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, and extract the local detail features of the target voxel in the region with the voxel as the unit and the step size set to 1;

[0117] S72: Calculate the color feature T C (v) in each target voxel through the mean calculated by the first moment and the variance calculated by the second moment;

[0118] S73: Use the LBP algorithm to extract the texture feature T LBP (v) in each target voxel for differentiating different drugs;

[0119] S74: Use the SURF algorithm to extract the feature points T SURF (v) in each target voxel;

[0120] S75: Concatenate the color feature, the texture feature, and the feature points to obtain the local detail feature in each target voxel:

[0121] T(v)=[T C (v); T LBP (v); T SURF (v)];

[0122] S8: Extract the foreground local detail feature and the foreground global topological invariance feature on the foreground image corresponding to each target voxel according to the disparity map of the foreground image;

[0123] S9: Update the truncation distance and weight of the TSDF model according to the original global topological invariance feature, the original local detail feature, the foreground local detail feature, and the foreground global topological invariance feature of each target voxel to obtain the TSDF global fusion model;

[0124] S91: Extract the original local detail feature T(v) and the original global topological invariance feature G(v) of each target voxel v according to the disparity image of the left-eye image and the right-eye image;

[0125] S92: Extract the foreground local detail feature T(u) and the foreground global topological invariance feature G(u) on the foreground image corresponding to each target voxel u according to the disparity map of the foreground image;

[0126] S93: Update the truncation distance diff(v) and weight w of the TSDF model i (v);

[0127] The truncation distance update formula is as follows:

[0128] diff(v) = |T(v) - T(u)| + |G(v) - G(u)|;

[0129] The weight update formula is as follows:

[0130] w i (v) = min(w i-1 (v) + λ i (v) + ξ(v), w μ );

[0131] where ξ(v) is the weight update value corresponding to the global topological invariance feature and local detail feature of each voxel; λ i (v) is the weight increment provided by the target voxel;

[0132] S94: Obtain the updated TSDF value according to the updated truncation distance and weight, which is expressed as:

[0133]

[0134] S95: Obtain the TSDF global fusion model according to the updated TSDF value and weight

[0135] S10: Use the marching cubes algorithm to calculate and obtain the isosurface composed of all target voxels with a weighted sum of directed distances of 0 in the TSDF global fusion model, and obtain the surface image of the target drug.

[0136] The real-time three-dimensional reconstruction method of the drug surface based on binocular vision according to the present invention uses the disparity image generated after binocular vision matching of the left-eye image and the right-eye image collected by the binocular camera as the region, and extracts the contour features, spatial structure features, and geometric shape features for distinguishing whether the drug is boxed, bottled, or bagged within the region, and concatenates them to form the global topological invariance features; extracts the color features, texture features, and pixel feature points for obtaining information such as the drug type, name, and manufacturer within the region, and concatenates them to form the local detail features; uses the global topological invariance features and local detail features of the corresponding voxels extracted from the disparity image and the foreground image of the target drug to update the truncation distance and weight in the TSDF model to obtain the TSDF global fusion model; in addition to the truncation distance value and weight for each voxel in the TSDF global fusion model, the global topological invariance features and local detail features of the voxel grid are also added. When calculating the isosurface with a weighted sum of directed distances equal to 0 using the TSDF global fusion model to obtain the surface of the target drug, a more real-time and accurate three-dimensional reconstruction of the drug surface is realized, providing a more accurate drug recognition model for the medicine-taking robot and improving the pose accuracy of grasping the drug.

[0137] Based on the above embodiments, the present invention further provides a real-time three-dimensional reconstruction system for the drug surface based on binocular vision, including:

[0138] A preprocessing module 100, configured to preprocess the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image;

[0139] A binocular vision stereo matching module 200, configured to perform binocular vision stereo matching on the left-eye image and the right-eye image to obtain a disparity image;

[0140] A target drug segmentation module 300, configured to detect the disparity image using an edge detection algorithm to distinguish the foreground image and the background image, and obtain the foreground image containing the target drug;

[0141] A target drug three-dimensional coordinate acquisition module 400, configured to obtain the three-dimensional coordinates of the corresponding points of the target drug according to the disparity map of the foreground image;

[0142] A TSDF model construction module 500, configured to project the disparity map of the foreground image into the three-dimensional space in reverse using the TSDF model according to the three-dimensional coordinates to obtain a three-dimensional feature volume, and use the voxels with TSDF values not greater than 1 in the three-dimensional feature volume as target voxels;

[0143] The feature extraction module 600 is configured to extract the original global topological invariance features and original local detail features of the target voxels according to the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image; and extract the foreground global topological invariance features and foreground local detail features of the target voxels according to the disparity map of the foreground image.

[0144] The TSDF global fusion model reconstruction module 700 is configured to update the truncation distance and weights according to the relevant features of each target voxel extracted by the feature extraction module, and obtain the TSDF global fusion model.

[0145] The drug surface acquisition module 800 is configured to calculate and obtain the isosurface formed by all target voxels with a weighted sum of directed distances of 0 in the TSDF global fusion model, and obtain the target drug surface image.

[0146] The real-time three-dimensional reconstruction system for drug surface based on binocular vision described in this embodiment is used to implement the aforementioned real-time three-dimensional reconstruction method for drug surface based on binocular vision. Therefore, the specific implementation manners in the real-time three-dimensional reconstruction system for drug surface based on binocular vision can be seen in the embodiment part of the aforementioned real-time three-dimensional reconstruction method for drug surface based on binocular vision. For example, the preprocessing module 100, the binocular vision stereo matching module 200, the target drug segmentation module 300, the target drug three-dimensional coordinate acquisition module 400, and the TSDF model construction module 500 are respectively used to implement steps S1, S2, S3, S4, and S5 in the aforementioned real-time three-dimensional reconstruction method for drug surface based on binocular vision; the feature extraction module 600 is used to implement steps S6, S7, and S8 in the aforementioned real-time three-dimensional reconstruction method for drug surface based on binocular vision; the TSDF global fusion model reconstruction module 700 and the drug surface acquisition module 800 are respectively used to implement steps S9 and S10 in the aforementioned real-time three-dimensional reconstruction method for drug surface based on binocular vision; therefore, the specific implementation manners can refer to the descriptions of the corresponding individual part embodiments and will not be elaborated here.

[0147] Specifically, based on the above embodiments, a boxed drug is selected, and the real-time three-dimensional reconstruction method for drug surface based on binocular vision provided by the embodiments of the present invention is used for three-dimensional reconstruction of the drug surface; referring to Figure 2 as shown, the modeling results of the boxed drug at different times are shown; specifically, based on the above embodiments, a bottled drug is selected, and the real-time three-dimensional reconstruction method for drug surface based on binocular vision provided by the embodiments of the present invention is used for three-dimensional reconstruction of the drug surface; referring to Figure 3 as shown, the modeling results of the bottled drug at different times are shown. According to Figure 2 and Figure 3It can be seen that as the modeling time increases, the details of the model can be better presented. Within 0.03 s, the target drug model has been clearly reconstructed, and the modeling results within a short time can still relatively completely reflect the geometric features of the target sample model.

[0148] Precision refers to the error between the three-dimensional model reconstructed for the drug and the true three-dimensional structure of the drug. To verify the modeling dimension precision of the reconstruction algorithm provided by the present invention, as shown in Table 1, the dimension errors between the three-dimensional models reconstructed for boxed drugs and bottled drugs and their true three-dimensional structures are given:

[0149] Table 1: Modeling Dimension Errors

[0150]

[0151] According to the data shown in Table 1, the absolute values of the modeling dimension errors of boxed drug 1 and bottled drug 2 do not exceed 0.05 cm; therefore, the modeling time of the method for real-time three-dimensional reconstruction of the drug surface based on binocular vision provided by the present invention is short and the modeling precision is high.

[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0153] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for the functions specified in one block or multiple blocks.

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements in the process Figure 1One process or multiple processes and / or boxes Figure 1 The functions specified in one box or multiple boxes.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 One process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one box or multiple boxes.

[0156] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A real-time three-dimensional reconstruction method for the surface of drugs based on binocular vision, characterized in that, Including: Preprocessing the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image, and performing binocular vision stereo matching to obtain a disparity image; Detecting the disparity image by using an edge detection algorithm, distinguishing the foreground image and the background image, obtaining the foreground image containing the target drug, and calculating the three-dimensional coordinates of the corresponding points of the target drug; According to the three-dimensional coordinates of the corresponding points of the target drug, using the TSDF algorithm to back-project the disparity map of the foreground image into the three-dimensional space to obtain a three-dimensional feature volume, and taking the voxels with TSDF values not greater than 1 in the three-dimensional feature volume as target voxels; According to the disparity image, extracting the contour features, spatial structure features and geometric shape features of the target voxels, and concatenating them to form the original global topological invariance features; According to the disparity image, extracting the color features, texture features and voxel feature points of the target voxels, and concatenating them to form the original local detail features; According to the disparity map of the foreground image, extracting the foreground local detail features and foreground global topological invariance features on the foreground image corresponding to the target voxels; According to the original global topological invariance features, the original local detail features, the foreground local detail features and the foreground global topological invariance features of each target voxel, updating the truncation distance and weight of the TSDF model to obtain the TSDF global fusion model; Calculating and obtaining the isosurface composed of all target voxels with a directed distance weighted sum of 0 in the TSDF global fusion model to obtain the surface image of the target drug.

2. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, wherein The preprocessing includes: Performing intensity transformation on the pixel points in the left-eye original image and the right-eye original image to obtain a new pixel point intensity value with variables k, a, b, and c to be optimized; taking each parameter as a particle in the particle swarm algorithm by taking values within a preset range according to a preset step size, and using the particle swarm algorithm update formula to update the velocity and position of the particle until the value of the fitness function converges to obtain the left-eye image and the right-eye image; The expression of the intensity transformation is: ; Among them, represents the coordinates of the pixel point in the original image, represents the image intensity value of this pixel point before preprocessing, represents the local mean, represents the global mean, represents the local standard deviation. The parameters to be optimized, k, a, b, and c, are all unknown parameters. Taking the local mean as the processing center, the original image is preprocessed; The particle swarm update formula is: , , Among them, represents the current iteration number, represents the current particle, represents the current dimension, represents the iteration number of times when the velocity of the particle, represents the iteration number of times when the position of the particle, represents the historical best solution of the current particle, represents the best solution of the global particle, is to adjust acceleration coefficient, is to adjust acceleration coefficient, and represent mutually independent random functions.

3. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, characterized in that, The performing binocular vision stereo matching to obtain a disparity image includes: Using the SURF algorithm to detect feature points in the left-eye image and the right-eye image, and performing sparse representation on the detected feature points; Calculating the gradient similarity between the sparse representation feature points of the left-eye image and the sparse representation feature points of the right-eye image as the initial cost of matching; Taking the global energy optimal strategy as the cost aggregation function, and using WTA to calculate the minimum cost aggregation to obtain the disparity value and obtain the disparity image.

4. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, wherein The using an edge detection algorithm to detect the disparity image, distinguishing the foreground image and the background image, and obtaining the foreground image containing the target drug includes: Establishing a consumption equation for different edges and obtaining the optimal solution of the consumption equation; Distinguishing the foreground image and the background image according to the optimal solution to obtain the foreground image containing the target drug.

5. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, characterized in that, The process of obtaining the three-dimensional coordinates of the corresponding points of the target drug includes: Take a point in space , is the focal length of the binocular camera, is the baseline, representing the distance between the optical centers of the binocular cameras; is the image point of point under the left-eye camera, is the image point of point under the right-eye camera, is the distance from to the imaging plane of the left-eye camera, is the distance from to the imaging plane of the right-eye camera, represents the position difference of the corresponding points of the point in the images formed by the binocular cameras, then the parallax is expressed as: ; Image points under the left-eye camera and image points under the right-eye camera The distance between them is expressed as: ; According to the similarity triangle theory, it can be obtained that: ; Parallax is in inverse proportion. If the point is not fixed, the positions of the image points on the left and right images during imaging will change with the movement of the point, and then the parallax will also change. Combining with the depth information of the third dimension, the three-dimensional coordinates of the point are expressed as: 。 6. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, characterized in that The obtaining of the original global topological invariance features includes: Taking the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, extract the global topological invariance features of the target voxels in the region; Use the snake contour feature extraction algorithm to obtain the edge contour features of the target drug within the region ; Extract features using Gabor transform to obtain the first feature vector , and obtain the second feature vector using GLCM gray histogram features , the spatial structure features of the target drug within the region , expressed as ; Use the Canny operator to detect and obtain the geometric shape features of the target drug within the region ; Concatenate the edge contour features, the spatial structure features, and the geometric shape features to obtain the global topological invariance features: 。 7. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, characterized in that The acquisition of the local detail features includes: Taking the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image as the region, taking voxels as the unit, setting the step size to 1, and extracting the local detail features of the target voxels in the region; Calculate the color features within each target voxel through the mean calculated by the first moment and the variance calculated by the second moment ; Extract the texture features used to distinguish different drugs within each target voxel using the LBP algorithm ; Extract feature points within each target voxel using the SURF algorithm ; Concatenate the color features, the texture features, and the feature points to obtain the local detail features within each target voxel: 。 8. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, characterized in that The acquisition of the TSDF global fusion model includes: Extract the target voxels according to the disparity image of the left-eye image and the right-eye image of the original local detail features and the original global topological invariance features ; Extract target voxels according to the disparity map of the foreground image Foreground local detail features on the corresponding foreground image And foreground global topological invariance features ; According to each target voxel's , , and , update the truncation distance and weight of the TSDF model; The truncation distance update formula is: ; The weight update formula is: ; Among them, is the weight update value corresponding to the global topological invariance feature and the local detail feature of the target voxel; is the weight increment provided for the target voxel; According to the updated truncation distance and weight, obtain the updated TSDF value, expressed as: ; Obtain the TSDF global fusion model based on the updated TSDF values and weights .

9. The real-time three-dimensional reconstruction method for the surface of a drug based on binocular vision according to claim 1, wherein Using the marching cubes algorithm in the TSDF global fusion model, calculate and obtain the isosurface composed of all target voxels with a directed distance weighted sum of 0.

10. A system based on the real-time three-dimensional reconstruction method of the drug surface based on binocular vision according to any one of claims 1 to 9, characterized in that, Including: A preprocessing module for preprocessing the left-eye original image and the right-eye original image collected by the binocular camera respectively to obtain the left-eye image and the right-eye image; A binocular vision stereo matching module for performing binocular vision stereo matching on the left-eye image and the right-eye image to obtain a disparity image; A target drug segmentation module for detecting the disparity image using an edge detection algorithm, distinguishing the foreground image and the background image, and obtaining the foreground image containing the target drug; A target drug three-dimensional coordinate acquisition module for obtaining the three-dimensional coordinates of the corresponding points of the target drug according to the disparity map of the foreground image; A TSDF model construction module for reversely projecting the disparity map of the foreground image into the three-dimensional space using the TSDF model according to the three-dimensional coordinates to obtain a three-dimensional feature volume, and taking the voxels with a TSDF value not greater than 1 in the three-dimensional feature volume as target voxels; A feature extraction module for extracting the original global topological invariance features and the original local detail features of the target voxels according to the disparity image generated by binocular vision stereo matching of the left-eye image and the right-eye image; extracting the foreground global topological invariance features and the foreground local detail features of the target voxels according to the disparity map of the foreground image; A TSDF global fusion model reconstruction module for updating the truncation distance and weight according to the relevant features of each target voxel extracted by the feature extraction module to obtain the TSDF global fusion model; A drug surface acquisition module for calculating and obtaining the isosurface composed of all target voxels with a directed distance weighted sum of 0 in the TSDF global fusion model to obtain the target drug surface image.