Robot stereoscopic vision depth estimation method, system, device, medium and product
By employing HSL space and improved Census cost calculation combined with adaptive smoothing filtering in the live-line operation robot, the depth estimation error problem in large-area weak texture and depth discontinuity regions was solved, and higher-precision depth information acquisition was achieved.
Patent Information
- Application Number
- CN202510965644.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
The depth estimation error of the live-line operation robot based on binocular stereo vision is relatively high in large areas with weak texture and discontinuous depth, which leads to safety hazards. Existing technologies are difficult to accurately obtain scene depth information.
We employ a method based on absolute difference calculation in HSL space, improved Census cost calculation, and adaptive smoothing guided filtering to fuse image information, thereby reducing depth estimation error and improving estimation accuracy.
It effectively reduces the depth estimation error in large areas with weak texture and depth discontinuity, and improves the accuracy and robustness of stereo vision depth estimation for live-line operation robots.
Smart Images

Figure CN120807606A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, and in particular to a robot stereo vision depth estimation method, system, device, medium and product. BACKGROUND
[0002] The precise operation of the live-line working robot for draining lines is an important part of ensuring the safe construction of the substation. In the operation process of the live-line working robot for draining lines, real-time acquisition of the depth information of the environment is the premise for autonomous and intelligent operation. At present, the live-line working robot for draining lines acquires depth information by carrying a laser radar, a monocular camera and a binocular stereo camera. The live-line working robot for draining lines based on the laser radar can obtain the depth information of a long-distance and large-range scene, but in a short-distance and large-area operation scene, the obtained depth information is too sparse, which can easily cause incomplete estimation of the scene depth information and bring safety hazards to the live-line operation of the live-line working robot for draining lines. The live-line working robot for draining lines based on monocular vision has the problems of complex algorithm, low estimation accuracy and the need for prior geometric knowledge of the scene. The live-line working robot for draining lines based on binocular stereo vision has the advantages of stable algorithm, dense depth results and high accuracy. Therefore, the live-line working robot for draining lines based on binocular stereo vision is the best choice for the live-line working robot for draining lines to obtain the scene depth in real time and accurately.
[0003] At present, the live-line working robot for draining lines based on binocular stereo vision mainly faces the difficulties of large-area weak texture and depth discontinuity in the operation scene, that is, the wall, road and other large-area texture single regions of the scene and the regions with depth faults such as deep pits and rocks. The depth estimation error of these regions is easy to mislead the movement and operation of the live-line working robot for draining lines in the scene, and the high depth estimation error is easy to cause safety hazards. SUMMARY
[0004] Therefore, the present application provides a robot stereo vision depth estimation method, system, device, medium and product, which solves the technical problem of high depth estimation error.
[0005] The first aspect of the present application provides a robot stereo vision depth estimation method, comprising:
[0006] acquiring stereo image information of an operation scene corresponding to a live-line working robot for draining lines; wherein the stereo image information comprises a left image and a right image;
[0007] calculating a first generation value of an absolute difference value based on an HSL space according to the left image and the right image;
[0008] calculate a second generation cost value of the left image and the right image based on the improved Census cost, and fuse the first generation cost value and the second generation cost value to obtain a fused generation cost value;
[0009] perform cost aggregation on the fused generation cost value based on adaptive smooth guided filtering to obtain a cost aggregated generation cost value;
[0010] perform cost calculation on the cost aggregated generation cost value to obtain a pixel disparity value of the left image and the right image;
[0011] perform disparity depth calculation on the work scene by using the pixel disparity value to obtain a depth estimation result of the work scene.
[0012] Preferably, the first generation cost value based on the absolute difference value in the HSL space is calculated according to the left image and the right image, comprising:
[0013] determine a generation cost value of an initial absolute difference value calculated in the HSL space of the pixel points of the left image and the right image according to the channel value of the pixel points of the left image and the channel value of the pixel points of the right image; wherein the generation cost value of the initial absolute difference value is:
[0014]
[0015] In the formula, H, S and L are hue channel, saturation channel and luminance channel respectively, p is a pixel to be matched of the left image, q is a pixel of the right input image, i is a channel index, and I i L (p) is a channel value corresponding to the pixel point p of the left image; I i R (q) represents a channel value corresponding to the pixel point q of the right image; C(p, q) is a generation cost value of the initial absolute difference value calculated in the HSL space of the pixel point p and the pixel point q;
[0016] normalize the generation cost value of the initial absolute difference value to determine the first generation cost value; wherein the first generation cost value is:
[0017]
[0018] In the formula, ρ(C AD (p, q), λ s ) is the first generation cost value, λ s is an adjustment parameter.
[0019] Preferably, the second generation cost value of the left image and the right image is calculated based on the improved Census cost, and the first generation cost value and the second generation cost value are fused to obtain the fused generation cost value, comprising:
[0020] obtaining a first matching cost of the input image based on a Census transform, wherein the input image comprises the left image and the right image, and the first matching cost is:
[0021]
[0022] wherein C census1 (p) is the first matching cost of pixel point p, I is the input image, (u, v) is the pixel coordinate, i, j are variables, m, n are the window length and width of the local neighborhood in which the pixel point p is located respectively, m', n' are the maximum integer not greater than half of m and n respectively, is the bitwise connection of bits, ξ is a comparison function, x, y are the horizontal and vertical coordinates of pixel point p respectively;
[0023] determining the local neighborhood window gray average value of the pixel point p and the gray average value of the entire input image according to the input image;
[0024] determining the second matching cost and the third matching cost of the pixel point p according to the local neighborhood window gray average value and the gray average value of the entire input image, wherein the second matching cost and the third matching cost are respectively:
[0025]
[0026]
[0027] wherein I p * is the local neighborhood window gray average value, I p is the gray value of pixel point p, I L is the gray value of the entire input image, I L * is the gray average value of the entire input image, C Census2 (p) is the second matching cost of pixel point p, C Census3 (p) is the third matching cost of pixel point p, ξ(I p , I L * is the Census value calculated by using the comparison function, is the Census value calculated by using the comparison function with the difference as the input;
[0028] bitwise connecting the first matching cost, the second matching cost and the third matching cost to obtain an initial second value;
[0029] normalizing the initial second value to obtain a second value;
[0030] Fusing the first generation cost value and the second generation cost value to obtain a fused generation cost value.
[0031] Preferably, the adaptive smoothing guided filtering is used to aggregate the fused generation cost value to obtain an aggregated generation cost value, including:
[0032] According to the coordinate information of the left image, a local adaptive smoothing parameter varying with image gradient in the left image is calculated;
[0033] Using the local adaptive smoothing parameter, a linear coefficient of adaptive smoothing guided filtering in each calculation local neighborhood window is calculated;
[0034] According to the linear coefficient, the fused generation cost value is aggregated to obtain the aggregated generation cost value.
[0035] Preferably, the aggregated generation cost value is used to calculate a cost to obtain a pixel disparity value of the left image and the right image, including:
[0036] The aggregated generation cost value is used to calculate a cost based on a winner-takes-all algorithm to obtain the pixel disparity value of the left image and the right image.
[0037] Preferably, the pixel disparity value is used to calculate a disparity depth of the work scene to obtain a depth estimation result of the work scene, including:
[0038] Based on a binocular stereo vision depth calculation formula, the pixel disparity value is used to calculate a disparity depth of the work scene to obtain a depth estimation result of the work scene; wherein the binocular stereo vision depth calculation formula is:
[0039]
[0040] In the formula, L is a depth estimation result of the work scene, F is a focal length of a binocular stereo camera corresponding to the obtained stereo image information, D is a pixel disparity value, and B is a baseline value of the binocular stereo camera corresponding to the obtained stereo image information.
[0041] In a second aspect, the present application further provides a robot binocular stereo vision depth estimation system, including:
[0042] An image acquisition module is configured to acquire stereo image information of a work scene corresponding to a live wire live working robot; wherein the stereo image information includes a left image and a right image;
[0043] A cost calculation module is configured to calculate a first generation cost value based on an absolute difference value in an HSL space according to the left image and the right image;
[0044] a cost fusion module configured to calculate a second generation cost of the left image and the right image based on an improved Census cost, and fuse the first generation cost and the second generation cost to obtain a fused generation cost;
[0045] a cost aggregation module configured to perform cost aggregation on the fused generation cost based on adaptive smooth guided filtering to obtain a cost aggregated cost;
[0046] a disparity calculation module configured to perform cost calculation on the cost aggregated cost to obtain a pixel disparity value of the left image and the right image;
[0047] a depth estimation module configured to perform disparity depth calculation on the working scene by using the pixel disparity value to obtain a depth estimation result of the working scene.
[0048] In a third aspect, the present application further provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor perform the steps of the robot stereo vision depth estimation method according to the first aspect.
[0049] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed to implement the steps of the robot stereo vision depth estimation method according to the first aspect.
[0050] In a fifth aspect, the present application further provides a computer program product, which comprises a computer program stored in a non-transitory computer readable storage medium, and the computer program comprises program instructions, wherein when the program instructions are executed by a computer, the computer performs the steps of the robot stereo vision depth estimation method according to the first aspect.
[0051] From the above technical scheme can be seen, the application obtains the luminance and texture feature information in the scene by the cost calculation method based on HSL space, calculates the first generation value of the absolute difference based on HSL space, and can effectively reduce the depth estimation error of the drainage line live working robot stereo matching method in a large area of weak texture region. The cost calculation method based on improved Census fuses the color similarity feature information and gradient feature information of the neighborhood pixels and the center pixel in the local neighborhood window, and takes into account the acquisition of the texture feature and depth discontinuity feature in the scene, which can effectively reduce the depth estimation error in the depth discontinuous region. The first generation value based on the absolute difference of HSL space and the second generation value based on the improved Census cost calculation are fused to obtain the fusion generation value, and the adaptive smoothing guided filtering is used for cost aggregation of the fusion generation value. The adaptive smoothing parameter of the image is added to the image according to the gradient, so as to solve the noise caused by the too sharp edge, and the depth estimation error in the depth discontinuous region can be further reduced. The cost calculation is performed on the cost aggregation value to obtain the pixel disparity value, the disparity depth of the working scene is calculated by using the pixel disparity value, and the depth estimation result of the working scene is obtained, so as to reduce the depth estimation error and improve the accuracy of the depth estimation. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0053] Figure 1 An application environment diagram of a robot stereo vision depth estimation method provided by the embodiment of the present application is shown in the figure.
[0054] Figure 2 A flowchart of a robot stereo vision depth estimation method provided by the embodiment of the present application is shown in the figure.
[0055] Figure 3 A structural schematic diagram of a robot stereo vision depth estimation system provided by the embodiment of the present application is shown in the figure.
[0056] Figure 4 A structural schematic diagram of an electronic device provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0057] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of the present application.
[0058] The robot stereo vision depth estimation method provided by the embodiments of the present application can be applied to the application environment as shown in Figure 1 The terminal 101 communicates with the server 102 through a network. The data storage system can store data required to be processed by the server 102. The data storage system can be integrated on the server 102, or placed on a cloud or other network server. The terminal 101 or the server 102 obtains stereo image information of a work scene corresponding to a live-wire live working robot; wherein the stereo image information includes a left image and a right image; a first generation value of an absolute difference value based on an HSL space is calculated according to the left image and the right image; a second generation value of the left image and the right image is calculated based on an improved Census cost, and the first generation value and the second generation value are fused to obtain a fused generation value; the fused generation value is aggregated based on adaptive smoothing guided filtering to obtain a generation value after cost aggregation; the generation value after cost aggregation is calculated to obtain a pixel disparity value of the left image and the right image; and the pixel disparity value is used for disparity depth calculation of the work scene to obtain a depth estimation result of the work scene.
[0059] The terminal 101 can be, but is not limited to, various personal computers, notebook computers, smart phones and tablet computers, etc.
[0060] The server 102 can be a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0061] As shown in Figure 2 The embodiments of the present application provide a robot stereo vision depth estimation method. Taking the terminal 101 or the server 102 in Figure 1 as an example, the method includes the following steps S1 to S6. Wherein:
[0062] Step S1, obtaining stereo image information of a work scene corresponding to a live-wire live working robot; wherein the stereo image information includes a left image and a right image.
[0063] The stereo image information of the work scene is obtained by a binocular stereo camera carried by the live-wire live working robot, including a left image and a right image.
[0064] It can be understood that the left image and the right image acquired by the binocular stereo camera are images of the same scene taken from two different perspectives. By calculating the difference between the two images, the depth information of the scene can be obtained.
[0065] Step S2, calculating a first generation value of an absolute difference value based on an HSL space according to the left image and the right image.
[0066] The HSL (Hue, Saturation, Lightness) space is a color space commonly used in image processing. Compared with the RGB color space, the HSL space is more consistent with human perception of color and has better robustness in processing color information.
[0067] In step S2, the hue (H), saturation (S), and lightness (L) channel values of each pixel point in the left image and the right image are first extracted. Then, for each matching pixel point p in the left image, a corresponding pixel point q is found in the right image, and the generation value of the absolute difference (AD) in the HSL space of the two pixel points is calculated. Considering the three dimensions of color information, it can more comprehensively reflect the difference between the two pixel points.
[0068] Step S3, calculating a second generation value of the left image and the right image based on the improved Census cost, and fusing the first generation value and the second generation value to obtain a fused generation value.
[0069] The improved Census cost is a cost calculation method considering the local neighborhood information of the pixel point. The traditional Census cost only considers the difference in gray value of the pixel point itself, while the improved Census cost combines the information of the pixel point and its local neighborhood, improving the accuracy and robustness of the matching.
[0070] Specifically, first, for each pixel point, the bit string representation of its local neighborhood is obtained based on Census transformation, and then the Census cost value of the corresponding pixel points in the left image and the right image is calculated.
[0071] In addition, the difference between the average gray value of the local neighborhood window and the average gray value of the entire input image is also combined, further enriching the information dimension of the cost calculation. By fusing the first generation value and the second generation value calculated by the improved Census cost, a more comprehensive and accurate fused generation value can be obtained.
[0072] Step S4, cost aggregation of the fused generation value based on adaptive smoothing guided filtering to obtain the generation value after cost aggregation.
[0073] The cost aggregation method of the adaptive smoothing guided filter is a cost aggregation method capable of effectively smoothing noise while preserving edge details. The method first calculates local adaptive smoothing parameters according to the coordinate information of the input image. These parameters change with the image gradient and can preserve sharp edge information in the edge region and perform smoothing processing in the non-edge region. Then, the linear coefficients of the adaptive smoothing guided filter in each local neighborhood window are calculated using these local adaptive smoothing parameters. These linear coefficients are used to weight the fusion cost value, thereby realizing cost aggregation. Through this method, the cost aggregation value can be obtained, which preserves the edge information of the image and effectively reduces the influence of noise, providing more accurate and reliable input for subsequent disparity calculation and depth estimation.
[0074] Step S5, performing cost calculation on the cost aggregation value to obtain the pixel disparity value of the left image and the right image.
[0075] In which, based on the winner-takes-all algorithm, the cost aggregation value is calculated to obtain the pixel disparity value of the left image and the right image.
[0076] In which, the winner-takes-all (WTA) algorithm is a commonly used disparity calculation method, and its basic idea is to select the disparity with the smallest cost value as the final disparity value. In step S5, for each pixel point, all possible disparity values are traversed, and the disparity with the smallest cost aggregation value is selected as the final disparity value of the pixel point. Through this method, the disparity value of each pixel point between the left image and the right image can be obtained, thereby constructing a disparity map.
[0077] Specifically, all possible disparity values are traversed, and for each pixel point, the cost aggregation value under different disparity values is compared, and the disparity with the smallest cost value is selected as the optimal disparity value of the pixel point. This process is accelerated through parallel calculation to improve processing efficiency. Through this step, the optimal disparity value of each pixel point between the left image and the right image is obtained, which constitutes a disparity map and provides key information for subsequent depth estimation.
[0078] Step S6, using the pixel disparity value to calculate the disparity depth of the work scene to obtain the depth estimation result of the work scene.
[0079] In which, based on the binocular stereo vision depth calculation formula, the depth information of the work scene can be accurately calculated through the pixel disparity value.
[0080] It should be noted that, by the cost calculation method based on the HSL space, the luminance and texture feature information in the scene is fully obtained, the first generation value based on the absolute difference value of the HSL space is calculated, the depth estimation error of the stereoscopic matching method of the live-wire charged operating robot in a large-area weak texture region can be effectively reduced, the color similarity feature information and gradient feature information of the neighborhood pixels and the center pixel in the local neighborhood window are fused based on the improved Census cost calculation method, the texture features and depth discontinuous features in the scene are considered, the depth estimation error of the depth discontinuous region can be effectively reduced, the fusion generation value is obtained by fusing the first generation value based on the absolute difference value of the HSL space and the second generation value based on the improved Census cost calculation, the cost aggregation is performed on the fusion generation value based on the adaptive smoothing guided filtering, the adaptive smoothing parameter of the image is added to the image according to the gradient, the noise caused by the too sharp edge can be solved, the depth estimation error of the depth discontinuous region can be further reduced, the pixel disparity value is obtained by performing the cost calculation on the cost aggregated generation value, the disparity depth of the operating scene is calculated by using the pixel disparity value, and the depth estimation result of the operating scene is obtained, so as to reduce the depth estimation error and improve the accuracy of the depth estimation.
[0081] In some embodiments, the first generation value based on the absolute difference value of the HSL space is calculated according to the left image and the right image, including:
[0082] In step S201, the initial absolute difference value generation value of the pixel points of the left image and the right image calculated in the HSL space is determined according to the channel value of the pixel point of the left image and the channel value of the pixel point of the right image; wherein the initial absolute difference value generation value is:
[0083]
[0084] In the formula, H, S, and L are the hue channel, the saturation channel, and the brightness channel respectively, p is the to-be-matched pixel of the left image, q is the pixel of the right input image, i is the channel index, and I i L (p) is the channel value corresponding to the pixel point p of the left image; I i R (q) represents the channel value corresponding to the pixel point q of the right image; C(p, q) is the initial absolute difference value generation value of the pixel point p and the pixel point q calculated in the HSL space;
[0085] In step S202, the initial absolute difference value generation value is normalized to determine the first generation value; wherein the first generation value is:
[0086]
[0087] In the formula, ρ(CAD (p, q),λ s ) is the first generation value, λ s is an adjustment parameter.
[0088] It can be understood that by normalizing the initial absolute difference value, the generation values between different channels can be comparable, and the adjustment parameter λ s can be used to control the strength of normalization to adapt to different application scenarios. Such processing helps to improve the accuracy and robustness of depth estimation.
[0089] In some embodiments, the second generation value of the left image and the right image is calculated based on the improved Census cost, and the first generation value and the second generation value are fused to obtain a fused generation value, comprising:
[0090] Step S301, for each input image, the first matching cost of the input image is obtained based on Census transformation; wherein the input image includes the left image and the right image; the first matching cost is:
[0091]
[0092] In the formula, C census1 (p) is the first matching cost of the pixel point p, I is the input image, (u, v) is the pixel coordinate, i, j are variables, m, n are the window length and width of the local neighborhood where the pixel point p is located respectively, m', n' are the maximum integers not greater than half of m and n, is the bitwise connection of the bit, ξ is the comparison function, x, y are the horizontal and vertical coordinates of the pixel point p respectively.
[0093] wherein the window size of the local neighborhood where the pixel point p(u, v) is located is set to m*n, and m and n are both odd numbers.
[0094] Step S302, according to the input image, the average gray value of the local neighborhood window where the pixel point p is located and the average gray value of the entire input image are determined.
[0095] In order to improve the robustness of the algorithm to weak texture areas and depth discontinuous areas, the color feature information and gradient feature information of the image are fused into the cost calculation process based on Census transformation. Specifically, the average gray value of the local neighborhood window where the pixel point p is located and the average gray value of the entire input image are calculated.
[0096] Wherein, the local neighborhood window gray average value refers to the arithmetic average value of the gray values of all pixel points in the local neighborhood window where the pixel point p is located. The gray average value of the whole input image refers to the arithmetic average value of the gray values of all pixel points in the whole input image. By calculating the difference between the local neighborhood window gray average value and the gray average value of the whole input image, the information dimension of the cost calculation can be further enriched, and the accuracy and robustness of the matching can be improved.
[0097] Step S303, determining the second matching cost and the third matching cost of the pixel point p according to the local neighborhood window gray average value and the gray average value of the whole input image; wherein, the second matching cost and the third matching cost are respectively:
[0098]
[0099]
[0100] In the formula, I p * is the local neighborhood window gray average value, I p is the gray value of the pixel point p, I L is the gray value of the whole input image, and I L * is the gray average value of the whole input image, C Census2 (p) is the second matching cost of the pixel point p, C Census3 (p) is the third matching cost of the pixel point p, ξ(I p *, I L * is the Census cost value calculated by using the comparison function, is the Census cost value calculated by using the comparison function with the input difference value;
[0101] Step S304, connecting the first matching cost, the second matching cost and the third matching cost by bits to obtain an initial second cost value.
[0102] Wherein, the first matching cost, the second matching cost and the third matching cost are connected by bits to obtain:
[0103]
[0104] In the formula, C census is the initial second cost value, that is, the cost value after the fusion of the color C ensus , the gradient C ensus and the traditional C ensus feature.
[0105] Step S305, normalizing the initial second cost value to obtain a second cost value.
[0106] Among them, the correlation between the left and right image pixels is calculated using the Hamming distance and normalized. The calculation formula is:
[0107]
[0108] Where C census , L (p) is the C of pixel p in the left image ensus Transformation result, C census,R (q) is the C of pixel q in the left image ensus Transformation result,λ c is the adjustment parameter, ρ(C census (p, q),λ c ) is the second generation value.
[0109] Step S306: Fuse the first generation value and the second generation value to obtain a fused generation value.
[0110] Among them, the cost value based on the absolute difference AD (Absolute Difference) in the HSL (Hue, Saturation, Lightness) space is combined with the cost value calculated based on the improved Census cost to obtain the final cost calculation value. The formula is:
[0111]
[0112] Where C(p, q) is the fusion cost.
[0113] It can be understood that the embodiment of the present application can fully utilize the color information and local neighborhood information in the image and improve the accuracy and robustness of matching by combining the first-generation value based on the absolute difference in HSL space and the second-generation value calculated based on the improved Census cost.
[0114] In some embodiments, performing cost aggregation on the fusion cost value based on the adaptive smoothing guided filtering to obtain the cost value after cost aggregation includes:
[0115] Step S401: Calculate a local adaptive smoothing parameter that changes with the image gradient in the input left image according to the coordinate information of the left image.
[0116] Locally adaptive smoothing parameters are dynamically adjusted based on image gradient information, aiming to strike a balance between edge preservation and noise smoothing. These parameters tend to maintain edge sharpness in image edges to avoid loss of edge information, while enhancing the smoothing effect in non-edge areas to suppress noise interference. This method allows for adaptive adjustment of smoothing strength, enabling the cost aggregation process to preserve important image features while effectively reducing the negative effects of noise.
[0117] The calculation process of the local adaptive smoothing parameter T is as follows:
[0118]
[0119] where i and j are the coordinates of the input left image, M(i, j) is the gradient value of the input left image, t and β are constant parameters set respectively, t is a truncation threshold, β is a constant term, and T(i, j) is the local adaptive smoothing parameter with the coordinates (i, j).
[0120] In step S402, the linear coefficients of the adaptive smoothing guided filter in each calculation local neighborhood window are calculated by using the local adaptive smoothing parameter.
[0121] where the linear coefficients a k and b k of the adaptive smoothing guided filter in each calculation local neighborhood window are calculated.
[0122]
[0123] where w is the local neighborhood window in the cost aggregation calculation, k is the serial number of the local neighborhood window, I is the guide image, i.e., the input left image, i is the pixel coordinate in the local neighborhood window w k , is the mean value of the pixel p in the neighborhood window w k , k and σ 2 k are the mean value and variance of the guide image I in the local neighborhood window w k ,
[0124] In step S403, the cost aggregation is performed on the fusion cost value according to the linear coefficients, to obtain the cost value after the cost aggregation.
[0125] where the cost aggregation is performed on the fusion cost value according to the linear coefficients, and the calculation formula is as follows:
[0126] ,
[0127] where C(p, q) is the fusion cost value, and C ASGF (p, q) is the cost value after the cost aggregation.
[0128] It can be understood that, by performing cost aggregation on the fusion cost value based on adaptive smooth guided filtering, the embodiments of the present application effectively fuse the edge information and local features of the image, while reducing the interference of noise. In this process, the introduction of the local adaptive smoothing parameter enables the filtering effect to be dynamically adjusted according to the image gradient information, both maintaining the sharpness of the edge and enhancing the smoothing effect. By calculating the linear coefficients of each local neighborhood window and using these linear coefficients to perform cost aggregation on the fusion cost value, the cost aggregation cost value is finally obtained. These cost values are not only more accurate and reliable.
[0129] In some embodiments, the depth estimation result of the work scene is obtained by using the pixel disparity value to calculate the disparity depth of the work scene.
[0130] Based on the binocular stereo vision depth calculation formula, the depth estimation result of the work scene is obtained by using the pixel disparity value to calculate the disparity depth of the work scene; wherein the binocular stereo vision depth calculation formula is:
[0131]
[0132] In the formula, L is the depth estimation result of the work scene, F is the focal length of the binocular stereo camera corresponding to the obtained stereo image information, D is the pixel disparity value, and B is the baseline value of the binocular stereo camera corresponding to the obtained stereo image information, i.e. the distance between the left camera and the right camera in the binocular stereo camera.
[0133] It can be understood that, by this formula, we can convert the calculated pixel disparity value into depth information of the work scene. In actual application, the focal length F and the baseline value B of the binocular stereo camera are usually known, so that accurate depth estimation results can be obtained by accurate disparity calculation. This process not only depends on high-quality image data and advanced stereo matching algorithms, but also needs to consider the complexity and diversity of the work scene to ensure the accuracy and robustness of the depth estimation. The embodiments of the present application effectively improve the performance of stereo vision depth estimation by combining the cost calculation based on HSL space, improved Census cost calculation, and adaptive smooth guided filtering, etc. technical means, and provide more reliable environmental perception ability for robot work.
[0134] Based on the same inventive concept, the embodiments of the present application also provide a robot stereo vision depth estimation system for implementing the robot stereo vision depth estimation method described above.
[0135] The implementation solution provided by the system to solve the problem is similar to the implementation solution described in the above method, and therefore the specific definitions in one or more robot stereo vision depth estimation system embodiments provided below can refer to the definitions of the robot stereo vision depth estimation method in the foregoing, which will not be described here again.
[0136] As Figure 3 shown, the embodiment of the present application provides a robot stereo vision depth estimation system, comprising:
[0137] An image acquisition module 100 is configured to acquire stereo image information of a work scene corresponding to a live wire work robot; wherein the stereo image information comprises a left image and a right image.
[0138] A cost calculation module 200 is configured to calculate a first generation cost value of an absolute difference value based on an HSL space according to the left image and the right image.
[0139] A cost fusion module 300 is configured to calculate a second generation cost value of the left image and the right image based on an improved Census cost, and fuse the first generation cost value and the second generation cost value to obtain a fused cost value.
[0140] A cost aggregation module 400 is configured to perform cost aggregation on the fused cost value based on adaptive smoothing guided filtering to obtain a cost aggregated cost value.
[0141] A disparity calculation module 500 is configured to perform cost calculation on the cost aggregated cost value to obtain a pixel disparity value of the left image and the right image.
[0142] A depth estimation module 600 is configured to perform disparity depth calculation on the work scene by using the pixel disparity value to obtain a depth estimation result of the work scene.
[0143] In some embodiments, the cost calculation module 200 is configured to:
[0144] Determine a first generation cost value of an initial absolute difference value calculated in an HSL space for pixel points of the left image and the right image according to channel values of the pixel points of the left image and channel values of the pixel points of the right image; wherein the first generation cost value of the initial absolute difference value is:
[0145]
[0146] In the formula, H, S, and L are hue channel, saturation channel, and luminance channel respectively, p is a pixel to be matched of the left image, q is a pixel of the right input image, i is a channel index, and I i L (p) is a channel value corresponding to the pixel point p of the left image; I i R(q) represents a channel value corresponding to a pixel point q of a right image; C(p, q) is an initial absolute difference value of pixel point p and pixel point q calculated in HSL space;
[0147] The initial absolute difference value is normalized to determine a first value; wherein the first value is:
[0148]
[0149] In the formula, ρ(C AD (p, q), λ s ) is the first value, and λ s is an adjustment parameter.
[0150] In some embodiments, the cost fusion module 300 is configured to:
[0151] For each input image, a first matching cost of the input image is obtained based on Census transformation; wherein the input image includes a left image and a right image; the first matching cost is:
[0152]
[0153] In the formula, C census1 (p) is the first matching cost of pixel point p, I is an input image, (u, v) is a pixel coordinate, i, j are variables, m, n are respectively a window length and a width of a local neighborhood window in which the pixel point p is located, m', n' are respectively the largest integer not greater than half of m and n, is a bitwise connection of a bit, ξ is a comparison function, and x, y are respectively a horizontal and vertical coordinate of the pixel point p;
[0154] According to the input image, a local neighborhood window gray average value of the pixel point p and a gray average value of the entire input image are determined;
[0155] According to the local neighborhood window gray average value and the gray average value of the entire input image, a second matching cost and a third matching cost of the pixel point p are determined; wherein the second matching cost and the third matching cost are respectively:
[0156]
[0157]
[0158] In the formula, I p * is the local neighborhood window gray average value, I p is a gray value of the pixel point p, I L is a gray value of the entire input image, and I L *C is the gray average value of the whole input image, C Census2 C is the second matching cost of pixel point p, C Census3 C is the third matching cost of pixel point p, C p C is the third matching cost of pixel point p, C L * C is the Census value calculated by using the comparison function, C is the Census value calculated by using the comparison function with the input difference value;
[0159] The first matching cost, the second matching cost and the third matching cost are connected by bits to obtain an initial second value;
[0160] The initial second value is normalized to obtain a second value;
[0161] The first value and the second value are fused to obtain a fused value.
[0162] In some embodiments, the cost aggregation module 400 is configured to:
[0163] The local adaptive smoothing parameter varying with the image gradient in the input left image is calculated according to the coordinate information of the left image;
[0164] The linear coefficient of the adaptive smoothing guided filter in each calculation local neighborhood window is calculated by using the local adaptive smoothing parameter;
[0165] The fused value is cost aggregated according to the linear coefficient to obtain a cost aggregated value.
[0166] In some embodiments, the disparity calculation module 500 is configured to:
[0167] The cost aggregated value is cost calculated based on the winner-takes-all algorithm to obtain the pixel disparity value of the left image and the right image.
[0168] In some embodiments, the depth estimation module 600 is configured to:
[0169] The disparity depth of the work scene is calculated based on the binocular stereo vision depth calculation formula by using the pixel disparity value to obtain the depth estimation result of the work scene; wherein the binocular stereo vision depth calculation formula is:
[0170]
[0171] In the formula, L is the depth estimation result of the work scene, F is the focal length of the binocular stereo camera corresponding to the obtained stereo image information, D is the pixel disparity value, and B is the baseline value of the binocular stereo camera corresponding to the obtained stereo image information.
[0172] As Figure 4As shown, the embodiment of the present application provides an electronic device, the electronic device 10 includes a memory 20 and a processor 30, the memory 20 stores a computer program, and the computer program is executed by the processor 30, so that the processor 30 executes the steps of the robot stereo vision depth estimation method in the above embodiment.
[0173] The embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed to realize the steps of the robot stereo vision depth estimation method in the above embodiment.
[0174] The embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer readable storage medium, and the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer executes the steps of the robot stereo vision depth estimation method in the above embodiment.
[0175] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, electronic device, computer storage medium and computer program product can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here.
[0176] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0177] It should be understood that, although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other sequences. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of the steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0178] In several embodiments provided by the present application, it should be understood that the disclosed system, electronic device, computer storage medium, computer program product and method can be implemented in other ways. For example, the above-described device embodiments are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or in other forms.
[0179] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0180] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0181] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for executing all or part of the steps of the method described in various embodiments of the present application by a computer device (which can be a personal computer, a server, or a network device, etc.). The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (English full name: Re-Only Memory, English abbreviation: ROM), a random access memory (English full name: Random Access Memory, English abbreviation: RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0182] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A robot stereo vision depth estimation method, characterized in that: include: Acquire stereoscopic image information of an operation scene corresponding to the drainage line live operation robot; wherein the stereoscopic image information includes a left image and a right image; Calculate a first generation value of an absolute difference based on the HSL space according to the left image and the right image; Calculating the second generation value of the left image and the right image based on the improved Census cost, and fusing the first generation value and the second generation value to obtain a fused cost value; performing cost aggregation on the fusion cost value based on an adaptive smoothing guided filter to obtain a cost value after cost aggregation; performing cost calculation on the cost value after cost aggregation to obtain a pixel disparity value between the left image and the right image; The disparity depth of the operation scene is calculated using the pixel disparity value to obtain a depth estimation result of the operation scene.
2. The robot stereo vision depth estimation method according to claim 1, characterized in that: The calculating of the first generation value of the absolute difference based on the HSL space according to the left image and the right image includes: Determine, based on the channel values of the pixels of the left image and the channel values of the pixels of the right image, a cost value of an initial absolute difference between the pixels of the left image and the right image calculated in the HSL space; wherein the cost value of the initial absolute difference is: Where H, S, L are hue, saturation and brightness channels respectively, p is the pixel to be matched in the left image, q is the pixel in the right input image, i is the channel index, and I i L (p) is the channel value corresponding to the pixel p in the left image; I i R (q) represents the channel value corresponding to pixel q in the right image; C(p, q) is the initial absolute difference cost value calculated in HSL space between pixel p and pixel q; Normalize the cost value of the initial absolute difference to determine the first cost value; wherein the first cost value is: Where, ρ(C AD (p, q),λ s ) is the first generation value, λ s is the adjustment parameter.
3. The robot stereo vision depth estimation method according to claim 1, characterized in that: The calculating the second-generation values of the left image and the right image based on the improved Census cost, and fusing the first-generation value and the second-generation value to obtain a fused cost value, includes: For each input image, a first matching cost of the input image is obtained based on a Census transform; wherein the input image includes the left image and the right image; and the first matching cost is: Where C census1 (p) is the first matching cost of pixel p, I is the input image, (u, v) is the pixel coordinate, i, j are variables, m, n are the window length and width of the local neighborhood where pixel p is located, m', n' are the maximum integers not greater than half of m and n respectively, is the bitwise concatenation of bits, ξ is the comparison function, x and y are the horizontal and vertical coordinates of pixel point p respectively; According to the input image, determining the grayscale average value of the local neighborhood window where the pixel point p is located and the grayscale average value of the entire input image; Determine the second matching cost and the third matching cost of the pixel point p based on the grayscale average of the local neighborhood window and the grayscale average of the entire input image; wherein the second matching cost and the third matching cost are respectively: Where, I p * is the average grayscale value of the local neighborhood window, I p is the gray value of pixel p, I L is the grayscale value of the entire input image, I L * is the grayscale average value of the entire input image, C Census2 (p) is the second matching cost of pixel p, C Census3 (p) is the third matching cost of pixel p, ξ(I p *, I L * ) is the Census cost value calculated using the comparison function, To calculate the Census cost value using the comparison function with the input as the difference; Concatenate the first matching cost, the second matching cost, and the third matching cost bit by bit to obtain an initial second generation value; Normalizing the initial second-generation value to obtain a second-generation value; The first generation value and the second generation value are fused to obtain the fused value.
4. The robot stereo vision depth estimation method according to claim 1, characterized in that: The performing cost aggregation on the fusion cost value based on the adaptive smoothing guided filtering to obtain the cost value after cost aggregation includes: Calculating a local adaptive smoothing parameter that changes with the image gradient in the input left image according to the coordinate information of the left image; Calculate the linear coefficient of the adaptive smoothing guided filter in each calculation local neighborhood window using the local adaptive smoothing parameter; Cost aggregation is performed on the fusion cost value according to the linear coefficient to obtain the cost value after the cost aggregation.
5. The robot stereo vision depth estimation method according to claim 1, characterized in that: The performing cost calculation on the cost value after the cost aggregation to obtain the pixel disparity value between the left image and the right image includes: Cost calculation is performed on the cost value after the cost aggregation based on a winner-takes-all algorithm to obtain the pixel disparity value between the left image and the right image.
6. The robot stereo vision depth estimation method according to claim 1, characterized in that: The calculating the disparity depth of the operation scene by using the pixel disparity value to obtain a depth estimation result of the operation scene includes: Based on a binocular stereoscopic vision depth calculation formula, the disparity depth of the work scene is calculated using the pixel disparity value to obtain a depth estimation result of the work scene; wherein the binocular stereoscopic vision depth calculation formula is: Where L is the depth estimation result of the working scene, F is the focal length of the binocular stereo camera corresponding to the stereo image information, D is the pixel disparity value, and B is the baseline value of the binocular stereo camera corresponding to the stereo image information.
7. A robot stereo vision depth estimation system, characterized in that: include: An image acquisition module is used to acquire stereoscopic image information of the working scene corresponding to the drainage line live working robot; wherein the stereoscopic image information includes a left image and a right image; A cost calculation module, configured to calculate a first generation value of an absolute difference based on an HSL space according to the left image and the right image; a cost fusion module, configured to calculate the second-generation value of the left image and the right image based on the improved Census cost, and fuse the first-generation value and the second-generation value to obtain a fused cost value; a cost aggregation module, configured to perform cost aggregation on the fusion cost value based on an adaptive smoothing guided filter to obtain a cost value after cost aggregation; a disparity calculation module, configured to perform cost calculation on the cost value after the cost aggregation to obtain a pixel disparity value between the left image and the right image; The depth estimation module is configured to calculate the parallax depth of the operation scene using the pixel parallax values to obtain a depth estimation result of the operation scene.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor performs the steps of the robot stereo vision depth estimation method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the robot stereo vision depth estimation method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to perform the steps of the robot stereo vision depth estimation method according to any one of claims 1 to 6.