Depth measurement method, binocular stereo vision system, device and storage medium
By multi-optimizing the original image pair of binocular stereo vision system, the problem of lack of depth information in specific structural scenes is solved, and the distance measurement accuracy and detection ability are improved.
Patent Information
- Application Number
- CN202210174111.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-02-24
AI Technical Summary
When binocular stereoscopic vision systems measure distances in specific structural scenarios, there is a problem of lack of depth information, especially when the optical axis is placed non-perpendicular to the ground, the parallax range is limited, resulting in a high failure rate of stereoscopic matching.
By performing the first and second optimization processes on the original image pairs taken by the binocular stereo vision system, the pixels in the horizontal direction of the image are reduced, the parallax influence is reduced, and the depth image is synthesized based on quality to improve the success rate of stereo matching.
It effectively improves the accuracy of ranging in low-altitude application scenarios, reduces the impact of parallax, and improves the success rate of stereo matching and detection capabilities.
Smart Images

Figure CN114519713B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and particularly to a depth measurement method, a binocular stereo vision system, a device, and a storage medium. Background Art
[0002] Currently, machine vision technologies have been widely applied in various fields, such as intelligent robots, intelligent industrial equipment, autonomous driving, etc. The ranging solutions in related technologies include, for example, laser ranging, ultrasonic ranging, time-of-flight (ToF) ranging, binocular stereo vision ranging, and radar ranging. Among them, some solutions can be further divided into active ranging and passive ranging.
[0003] Binocular stereo vision is a method for obtaining three-dimensional geometric information of an object from multiple images based on the parallax principle. A binocular stereo vision system generally obtains two digital images of a measured object from different angles by two cameras, and restores the three-dimensional geometric information of the object based on the parallax principle to reconstruct the three-dimensional contour and position of the object. It has the advantages of high efficiency, appropriate accuracy, simple system structure, and low cost.
[0004] The binocular stereo vision ranging solution can be further divided into active ranging and passive ranging specifically.
[0005] As Figure 1A shown, a schematic diagram of the principle of an active ranging solution based on binocular stereo vision is schematically shown. In this solution, an infrared light projector 101 projects a preset spot pattern on the measured object. For example, a relatively dense and uniform grating is projected onto the measured object, and then the left and right cameras 102 and 103 respectively capture images of the measured object, and the height difference and depth information between the measured points of the measured object 104 are calculated in combination with the geometric relationship between the left and right cameras. Its advantages are simple calculation and relatively high measurement accuracy, and precise measurement can be performed on flat surface areas without obvious texture and shape changes.
[0006] As Figure 1B shown, a schematic diagram of the principle of a passive ranging solution based on binocular stereo vision is schematically shown. Different from Figure 1A the above, the infrared projector is omitted, the left and right cameras 102 and 103 capture images, and the optical triangle principle is used to determine the depth information of each point of the object 104 in space according to the geometric relationship. The passive binocular stereo vision solution is often used for ranging of rough or objects with significant texture information.
[0007] Whether it is active ranging or passive ranging, stereo matching needs to be performed before calculating depth information. Specifically, two images captured by a binocular camera at their respective viewpoints form an image pair. By detecting the image blocks of the same object in the two images of the image pair and performing stereo matching, a pair of pixel points corresponding to the same physical point in the two images is determined, and the disparity between the pixel point pairs is obtained. Then, ranging calculations are performed based on the disparity map to obtain depth information. The depth information of each pixel point constitutes a depth image, and the depth image can be further transformed into the coordinates of each point in the world coordinate system to obtain a point cloud map.
[0008] However, binocular stereo vision technology has certain problems in scenes with specific structures. A common usage scenario is that a binocular stereo vision system is placed on the ground at a relatively low height. Whether it is active or passive binocular stereo vision ranging, when it approaches an object placed on the ground non-vertically along the optical axis direction to measure depth information, its two cameras will be restricted by the geometric structure. The lower the placement height of the two cameras, the greater the disparity between the captured images. The disparity range of stereo matching is limited, which will result in an increased failure rate of stereo matching and a large number of pixels lacking depth information in the depth image.
[0009] In addition, although time-of-flight (ToF) ranging using lasers is widely used in applications such as floor cleaning robots for stereo depth measurement close to the ground, the cost of lidar is much higher than that of binocular stereo vision systems and it is impossible to replace binocular stereo vision systems.
[0010] Therefore, how to find a binocular stereo vision system that can improve the ranging accuracy in low-height application scenarios has become an urgent technical problem in the industry.
[0011] Inventive News
[0012] In view of the above-mentioned shortcomings of related technologies, the purpose of the present disclosure is to provide a depth measurement method, a binocular stereo vision system, a device, and a storage medium to solve the problem of missing depth information in ranging of a binocular stereo vision system in a specific structure scene in related technologies.
[0013] The first aspect of the present disclosure provides a depth measurement method, which is applied to a binocular stereo vision system. The depth measurement method includes: acquiring an original image pair captured by the binocular stereo vision system; respectively performing a first optimization process and a second optimization process on the original image pair to obtain a first optimized image pair and a second optimized image pair. Wherein, the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image. The processed image is each original image in the original image pair; performing stereo matching based on the first optimized image pair and calculating to obtain a first depth map; and performing stereo matching based on the second optimized image pair and calculating to obtain a second depth image; synthesizing the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image.
[0014] In an embodiment of the first aspect, the first optimization process and the second optimization process are performed based on at least one adjustable image optimization factor, so as to adjust the disparity between the paired corresponding pixels in each of the first optimized image pair and the second optimized image pair.
[0015] In an embodiment of the first aspect, the image optimization factor is a preset image optimization factor, which is used to control the pixel reduction effects of the first optimization process and the second optimization process in an opposite strength manner.
[0016] In an embodiment of the first aspect, the first optimization process and the second optimization process are respectively image processing methods of selectively discarding and selectively retaining some pixels of the processed image in the horizontal direction of the image. The first quantity of the part of the pixels selectively discarded by the first optimization process is determined by a first image optimization factor. The second quantity of the part of the pixels selectively retained by the second optimization process is determined by a second image optimization factor. The first image optimization factor and the second image optimization factor are the same or different.
[0017] In an embodiment of the first aspect, the first image optimization factor and the second image optimization factor are the same preset image optimization factor.
[0018] In an embodiment of the first aspect, the preset image optimization factor is used to control the pixel discarding effect of the first optimization process and the pixel retaining effect of the second optimization process in an opposite strength manner.
[0019] In an embodiment of the first aspect, the preset image optimization factor is positively correlated with one of the first quantity and the second quantity, and negatively correlated with the other.
[0020] In an embodiment of the first aspect, before synthesizing the first depth image and the second depth image, it further includes: restoring the first depth image and / or the second depth image to the original image resolution by interpolation.
[0021] In an embodiment of the first aspect, the first optimization process includes: cropping and discarding an image area from the processed image to obtain an optimized image for stereo matching; wherein, the vertical image resolution between the cropped image area and the processed image is the same, and the ratio of the horizontal image resolution of the cropped image area to the processed image is 1:N, where N is an integer greater than 1; and / or, the second optimization process includes: sampling a pixel column at equal intervals from the processed image along the horizontal direction of the image and discarding the image area between each pixel column to obtain an optimized image for stereo matching; wherein, the pixel column has the same vertical image resolution as the processed image, and the ratio of the interval to the width of the processed image in the horizontal direction of the image is 1:P, where P is an integer greater than 1.
[0022] In an embodiment of the first aspect, both N and P are image optimization factors M.
[0023] In an embodiment of the first aspect, the cropping and discarding an image area from the processed image includes: centrally cropping and discarding the image area from the processed image.
[0024] In an embodiment of the first aspect, before synthesizing the first depth image and the second depth image, it further includes: performing image interpolation based on the second depth image to restore it to the original image resolution.
[0025] In an embodiment of the first aspect, synthesizing the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image includes: for each pixel point of the object point in the target depth image, selecting the corresponding position pixel point with better quality in the first depth image or the second depth image for filling.
[0026] In an embodiment of the first aspect, the binocular stereo vision system is set such that its optical axis direction is not perpendicular to a carrying surface; wherein, the carrier of the binocular stereo vision system is movably disposed on the carrying surface.
[0027] In an embodiment of the first aspect, the binocular stereo vision system is set to be close to the carrying surface, and the distance between the binocular stereo vision system and the carrying surface is positively correlated with the detection distance of the binocular stereo vision system.
[0028] The second aspect of the present disclosure provides a depth measurement device applied to a binocular stereo vision system. The depth measurement device includes: an image acquisition module for acquiring an original image pair captured by the binocular stereo vision system; a multi-channel optimization module for respectively performing a first optimization process and a second optimization process on the original image pair to obtain a first optimized image pair and a second optimized image pair, where the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image, and the processed image is each original image in the original image pair; a multi-channel depth calculation module for performing stereo matching based on the first optimized image pair and calculating to obtain a first depth map, and performing stereo matching based on the second optimized image pair and calculating to obtain a second depth image; and a synthesis module for synthesizing the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image.
[0029] The third aspect of the present disclosure provides a binocular stereo vision system, including: one or a pair of camera units for collecting an original image pair with parallax; a processing unit communicatively connected to each camera unit for running program instructions to execute the depth measurement method according to any one of the first aspect.
[0030] The fourth aspect of the present disclosure provides a motion device, including: a motion device body movably disposed on a bearing surface; the binocular stereo vision system according to the third aspect disposed on the motion device body and configured such that its optical axis direction is not perpendicular to the bearing surface.
[0031] In an embodiment of the third aspect, the binocular stereo vision system is configured such that its optical axis direction corresponds to the forward, backward, or lateral direction of the motion device body.
[0032] In an embodiment of the third aspect, the bearing surface is the ground, and the motion device includes a mobile robot.
[0033] The fifth aspect of the present disclosure provides a computer device, including: a communicator, a memory, and a processor; the communicator is used for external communication; the memory stores program instructions; the processor is used for running the program instructions to execute the depth measurement method according to any one of the first aspect.
[0034] The sixth aspect of the present disclosure provides a computer-readable storage medium storing program instructions, and the program instructions are run to execute the depth measurement method according to any one of the first aspect.
[0035] As described above, the embodiments of the present disclosure provide a depth measurement method, a binocular stereo vision system, a device, and a storage medium. The method includes obtaining an original image pair captured by the binocular stereo vision system; respectively performing a first optimization process and a second optimization process on the original image pair to obtain a first optimized image pair and a second optimized image pair, where the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image, and the processed image is each original image in the original image pair; performing stereo matching on the first optimized image pair and calculating to obtain a first depth map; and performing stereo matching on the second optimized image pair and calculating to obtain a second depth image; synthesizing the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image. By performing multi-path optimization processing on the original image pair captured by the binocular stereo system to reduce the reduced parallax and then performing stereo matching to obtain the depth image, and then selectively synthesizing the target depth image according to the quality, the influence of parallax is reduced, the success rate of stereo matching is improved, and the detection ability of the binocular stereo vision system is improved. Description of the Drawings
[0036] Figure 1A Show a schematic diagram of the principle of an active ranging scheme based on binocular stereo vision.
[0037] Figure 1B Show a schematic diagram of the principle of a passive ranging scheme based on binocular stereo vision.
[0038] Figure 2 Show a schematic diagram of the principle of parallax calculation in an example.
[0039] Figure 3 Show a schematic flowchart of the depth measurement method in an embodiment of the present disclosure.
[0040] Figure 4A Show a schematic diagram of the implementation principle of the first optimization process in a specific example of the present disclosure.
[0041] Figure 4B Show a schematic diagram of the implementation principle of the second optimization process in a specific example of the present disclosure.
[0042] Figure 5A Show a schematic flowchart of the depth measurement method in an application example of the present disclosure.
[0043] Figure 5B Show a schematic flowchart of the first optimization process, stereo matching, and depth calculation in an application example of the present disclosure.
[0044] Figure 5C Show a schematic flowchart of the second optimization process, stereo matching, and depth calculation in an application example of the present disclosure.
[0045] Figure 6A and Figure 6B Show a schematic diagram of image comparison before and after implementing the depth measurement method in an application example.
[0046] Figure 7 Show a schematic diagram of the module of a depth measurement device in an embodiment of the present disclosure.
[0047] Figure 8 Show a schematic diagram of the structure of a binocular vision system in an embodiment of the present disclosure.
[0048] Figure 9 Show a schematic diagram of the structure of a computer device in an embodiment of the present disclosure. Detailed implementation manners
[0049] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the information disclosed in the present disclosure. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in the present disclosure can also be modified or changed according to different viewpoints and application systems without departing from the spirit of the present disclosure. It should be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0050] The following takes the accompanying drawings as a reference and details the embodiments of the present disclosure so that those skilled in the technical field to which the present disclosure belongs can easily implement it. The present disclosure can be embodied in many different forms and is not limited to the embodiments described herein.
[0051] In the description of the present disclosure, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc., mean that the specific features, structures, materials, or characteristics represented in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics represented can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples represented in the present disclosure and the features of the different embodiments or examples.
[0052] In addition, the terms "first" and "second" are only used for the purpose of indication and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means two or more unless otherwise specifically defined.
[0053] In order to clearly illustrate the present disclosure, devices irrelevant to the description are omitted, and the same or similar components throughout the specification are given the same reference numerals.
[0054] Throughout the specification, when it is said that a certain device is "connected" to another device, this includes not only the case of "direct connection", but also the case of "indirect connection" with other elements placed therebetween. In addition, when it is said that a certain device "includes" a certain component, unless there is a particularly contrary record, it does not exclude other components, but means that other components may also be included.
[0055] Although in some instances the terms first, second, etc. are used herein to denote various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first interface and the second interface, etc. are indicated. Furthermore, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprising", "including" indicate the presence of the stated features, steps, operations, elements, modules, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, steps, operations, elements, modules, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, or meaning any one or any combination. Thus, "A, B or C" or "A, B and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A, B and C". An exception to this definition only occurs when the combination of elements, functions, steps or operations are mutually exclusive in some way.
[0056] The technical terms used herein are only for referring to specific embodiments and are not intended to limit the present disclosure. The singular forms used herein also include the plural forms as long as the statement does not clearly indicate the contrary meaning. The meaning of "including" used in the specification is to embody specific characteristics, regions, integers, steps, operations, elements and / or components, and does not exclude the existence or addition of other characteristics, regions, integers, steps, operations, elements and / or components.
[0057] Relative spatial terms such as "lower", "upper", etc. may be used to more easily illustrate the relationship of one device relative to another device illustrated in the drawings. Such terms refer to not only the meaning indicated in the drawings, but also other meanings or operations of the device in use. For example, if the device in the drawing is flipped, a certain device that was previously described as "lower" than other devices is then described as "upper" than other devices. Therefore, the exemplary term "lower" includes both above and below. The device may be rotated 90° or other angles, and the relative spatial terms are also interpreted accordingly.
[0058] Although not defined differently, including the technical and scientific terms used herein, all terms shall have the same meaning as generally understood by those skilled in the technical field to which this disclosure pertains. Terms defined in commonly used dictionaries are additionally interpreted to have meanings consistent with the relevant technical literature and the currently presented information. As long as they are not defined, they shall not be over-interpreted as ideal or overly formulaic meanings.
[0059] As Figure 2 shown, a schematic diagram of the principle of a binocular stereo vision system calculating depth information using parallax in an example is presented.
[0060] Exemplarily, the binocular stereo vision system may include a pair of imaging units, namely a first imaging unit and a second imaging unit. Each of the imaging units may be implemented by a camera. The first imaging unit and the second imaging unit have the same configuration, such as having the same focal length f, acquiring images at the same image resolution (Image Resolution, that is, the number of pixels in the image, such as 1024×768), etc., and the width of the image in the horizontal direction is set as L. The first imaging unit and the second imaging unit may be arranged side by side on the left and right, and their optical axis directions may be parallel to each other or intersect at the same inclination angle. Figure 2 shows that the optical axis directions of the first imaging unit and the second imaging unit are parallel to each other. The optical centers O L and O R of the connection line is called the baseline, and the baseline length is b.
[0061] The first imaging unit and the second imaging unit synchronously capture images of a scene, which can be called the first original image and the second original image, having the same resolution. Let the width of their images in the horizontal direction be L; the first original image and the second original image form an original image pair. There will be points in the first original image and the second original image that correspond to the same object in the captured image of the scene, which can be called object points. Figure 2 schematically shows the object point A. Due to the different perspectives of the first imaging unit and the second imaging unit, the positions of the same object point A imaged at the corresponding homologous pixel points A1 and A2 in the first and second original images are different in the horizontal direction of the image. For example, taking the horizontal direction of the image as the X-axis and the vertical direction of the image as the Y-axis, then the X coordinates of the corresponding homologous pixel points A1 and A2 of the same object point in the first original image and the second original image are different, presenting as X in the first original image shown in Figure 1 L and X in the second original image R . The parallax is generated from this coordinate difference in the horizontal direction of the image and is expressed as d = |X L - X R |.
[0062] The depth information to be calculated is Z. According to the principle of similar triangles in geometric relations, Z / b = (Z - f) / A1A2, and A1A2 = b - (X L - L / 2) - (L / 2 - X R ) = b - (X L - X R ) Then it can be calculated that Z / b = (Z - f) / A1A2 = (Z - f) / (b - (X L - X R )) After rearrangement, we can get Z = b×f / (X L - X R ).
[0063] Stereo matching is to match each pair of corresponding pixels. In related technologies, binocular stereo matching can be divided into four steps: matching cost calculation, cost aggregation, disparity calculation, and disparity optimization. Matching cost calculation The purpose of matching cost calculation is to measure the correlation between the pixels to be matched and the candidate pixels. Whether two pixels are the same object point or not, the matching cost can be calculated through the matching cost function. The smaller the cost, the greater the correlation and the greater the probability that they are the same object point. Before searching for the same object point for each pixel, a disparity search range D (Dmin~Dmax) is often specified. When searching for the disparity, the range is limited within D, and a three-dimensional matrix C with a size of W×H×D (W is the image width, H is the image height) is used to store the matching cost values of each pixel at each disparity within the disparity range. The matrix C is usually called DSI (Disparity Space Image). There are many methods for calculating the matching cost. In traditional photogrammetry, methods such as absolute difference of gray values (AD, Absolute Differences), sum of absolute differences of gray values (SAD, Sum of Absolute Differences), and normalized cross-correlation coefficient (NCC, Normalized Cross-correlation) are used to calculate the matching cost between two pixels; in computer vision, methods such as mutual information (MI, Mutual Information) method, Census transform (CT, Census Transform) method, Rank transform (RT, Rank Transform) method, and BT (Birchfield and Tomasi) method are often used as the calculation methods for the matching cost.
[0064] The purpose of cost aggregation is to enable the cost value to accurately reflect the correlation between pixels. The calculation of the matching cost in the previous step often only considers local information, and calculates the cost value based on the pixel information within a certain-sized window in the neighborhoods of two pixels. This is easily affected by image noise. Cost aggregation is to establish the connection between adjacent pixels and optimize the cost matrix according to certain criteria, such as adjacent pixels should have continuous disparity values, and then propagate it to the areas with low signal-to-noise ratio and poor matching effects. Commonly used cost aggregation methods include the scan line method, the dynamic programming method, the path aggregation method in the SGM algorithm, etc.
[0065] Disparity calculation is to determine the optimal disparity value for each pixel through the cost matrix after cost aggregation. Usually, the Winner-Takes-All (WTA) algorithm is used for calculation. The cost aggregation step is a crucial step in stereo matching and directly determines the accuracy of the algorithm.
[0066] The purpose of disparity optimization is to further optimize the disparity map obtained in the previous step and improve the quality of the disparity map, including steps such as eliminating incorrect disparities, appropriate smoothing, and sub-pixel accuracy optimization. For example, in related technologies, the Left-Right Check algorithm is used to eliminate incorrect disparities caused by occlusion and noise; the algorithm for eliminating small connected regions is used to eliminate isolated outliers; smoothing algorithms such as Median Filter and BilateralFilter are used to smooth the disparity map; in addition, some methods that can effectively improve the quality of the disparity map, such as Robust Plane Fitting, Intensity Consistent, and Locally Consistent, are also often used.
[0067] From the above, it can be seen that the size of the disparity is closely related to the accurate effect of stereo matching. In the calculation of depth information, according to Z = b × f / (X L -X R)It can be seen that there is a negative correlation between depth and parallax. Common robots moving on the ground, such as sweeping robots, need to set the binocular stereo vision system at a height extremely close to the ground (for example, less than 20 cm). Moreover, the optical axis direction of the binocular stereo vision system (obtained according to the optical axis equivalence of the binocular cameras) is not set perpendicular to the ground, that is, not vertically upward or downward. For example, it can be set forward, backward, laterally, etc. The closer to the ground, the depth information of the object points in the ground relative to the binocular vision system will decrease, and the corresponding parallax will increase. As mentioned before, the parallax during stereo matching needs to be within a preset range. Therefore, too large a parallax may make it difficult to perform stereo matching between corresponding pixels, resulting in problems such as matching errors or missing depth information. Although there are currently technologies that place the binocular cameras closer to help improve stereo matching in applications with a low-altitude ground plane view, this will lead to a shortening of the detection distance of the binocular stereo vision system, that is, it is necessary to sacrifice the ranging ability, and the application is limited.
[0068] In view of this, in the embodiments of the present disclosure, a solution is provided. By multiple different optimization processing methods for reducing parallax, the original image pair is processed separately, and depth maps of each path are obtained through stereo matching respectively, and then merged based on quality to finally obtain the target depth map, so as to address the problem that the parallax of the same object point exceeds the range in scenarios such as when the optical axis is set low and not perpendicular to the ground, resulting in stereo matching failure, and effectively improve the quality of the obtained depth map.
[0069] As Figure 3 shown, a flowchart of the depth measurement method in the embodiments of the present disclosure is shown. The depth measurement method is applied to a binocular stereo vision system.
[0070] In Figure 3 the embodiment, the depth measurement method includes:
[0071] Step S301: Obtain the original image pair captured by the binocular stereo vision system.
[0072] In some embodiments, in the case of using two imaging units, the original image pair is obtained according to the outputs after the first imaging unit and the second imaging unit are synchronously photographed. Alternatively, in the case of using one imaging unit, the original image pair can be obtained by taking pictures with one imaging unit successively in two perspectives.
[0073] Step S302: Perform a first optimization process and a second optimization process on the original image pair respectively to obtain a first optimized image pair and a second optimized image pair.
[0074] Among them, the image to be processed is each original image in the original image pair, namely the first original image and the second original image. The first optimization process and the second optimization process are different image processing methods for reducing the pixels of the image to be processed in the horizontal direction of the image. Specifically, the first optimization process is respectively performed on the first original image and the second original image in the original image pair to obtain two optimized images to form a first optimized image pair; and, the second optimization process is respectively performed on the first original image and the second original image in the original image pair to obtain two optimized images to form a second optimized image pair. By reducing the original pixels of the first original image and the second original image in the horizontal direction of the image, the influence of parallax is reduced.
[0075] In a possible example, the method for reducing the pixels of the image to be processed in the horizontal direction of the image, for example, includes: selecting an image area and discarding it; selecting an image area to retain and discarding the remaining image area, etc. For example, cropping and discarding a part of the image area in the horizontal direction of the image, or discarding the remaining image area after sampling pixels along the horizontal direction of the image, etc.
[0076] In some embodiments, the first optimization process and the second optimization process are performed based on at least one adjustable image optimization factor, so as to adjust the parallax between the paired corresponding pixel points in the two optimized images included in the first optimized processed image and / or the second optimized processed image respectively.
[0077] In a possible example, the first optimization process may be to select and discard an image area, and the first quantity of the partial pixels selected and discarded by the first optimization process is determined by a first image optimization factor, and the image optimization factor may be a coefficient related to the size of the selected and discarded image area. For example, the size of the original image is a multiple of the size of the selected and discarded image area; and / or, the number of pixels in the horizontal direction and / or the vertical direction of the selected and discarded image area, etc., that is, a coefficient related to the image resolution, or may also be a coefficient related to the physical size of the horizontal direction and / or the vertical direction of the selected and discarded image area, etc.
[0078] In a possible example, the second optimization process may be to select and retain an image area in the original image and discard another part of the image area. The way of selecting and retaining the image area is, for example, to perform operations such as pixel sampling along the horizontal direction of the image. The quantity of the partial pixels selected and retained by the second optimization process can be determined by a second image optimization factor, and the second image optimization factor may be a coefficient related to the size of the selected and retained image area. For example, the coefficient of the sampling interval; and / or, the number of pixels in the horizontal direction and / or the vertical direction of the image area sampled each time, etc.
[0079] It is understandable that in a scene where the binocular stereo vision system is set close to the ground, the parallax increases as the height of the binocular stereo vision system decreases. By taking an appropriate value of the image optimization factor, the parallax between each pair of pixels with the same name in the first optimized image pair and the second optimized image pair obtained after the optimization process can fall within a preset parallax range suitable for stereo matching. Exemplarily, the value of the image optimization factor can be set according to factors such as different camera units with different configuration parameters, different optical structure types (such as active binocular or passive binocular), and actual usage scenarios (such as low altitude) to meet the parallax adjustment requirements in different practical application scenarios.
[0080] In some embodiments, the first image optimization factor and the second image optimization factor may be different or the same.
[0081] like Figure 4A and Figure 4B , respectively demonstrating the implementation principles of the first optimization process and the second optimization process through specific examples.
[0082] You can refer to Figure 4A As shown, the first optimization process may exemplarily include: cropping and discarding an image area 402 from a processed image 401 (represented by a cross shading) to obtain an optimized image for stereo matching, wherein the processed image is the first original image and the second original image. The cropped image area 402 (black shading) may be the same as the vertical resolution of the image between the processed images, and the ratio of the cropped image area to the horizontal resolution of the image of the processed image is 1:N, where N is an integer greater than 1. Optionally, the cropped image area may be obtained by cropping in the center of the processed image in the horizontal direction. Figure 4A In this example, the processed image is assumed to have a resolution of X×Y, that is, the image contains X pixels in the horizontal direction and Y pixels in the vertical direction. After the processed image is cropped in the center of the image horizontally (X / N)×Y, the remaining part with a resolution of (X×(N-1) / N)×Y is obtained as the processed optimized image.
[0083] For example, you can refer to Figure 4B As shown, a schematic diagram showing the principle of the second optimization process in an example is shown. Exemplarily, the second optimization process may include: sampling a pixel column 403 at equal intervals from the processed image 401 along the horizontal direction of the image, discarding the image area between the pixel columns 403, so as to obtain an optimized image for stereo matching. The pixel column has the same vertical resolution as the processed image, and the ratio of the width of the interval to the processed image in the horizontal direction of the image is 1:P, where P is an integer greater than 1. Figure 4BIn the example, a second optimization process is performed on the processed image of X×Y. Sampling is performed at intervals of X / P along the horizontal direction of the image, and n pixel columns (i.e., n×Y, where n is an integer greater than 1) are sampled each time. In the figure, n = 1 is set for easy calculation. The image regions between the pixel columns are discarded, and then P pixel columns are obtained, forming an optimized image with an image resolution of P×Y.
[0084] N and P can be set independently or can be set through an image optimization factor. For example, both N and P are the preset image optimization factor M.
[0085] Therefore, it can be understood that the image optimization factor can be a preset image optimization factor and can be used to control the pixel reduction effects of the first optimization process and the second optimization process in the opposite strength.
[0086] Furthermore, the preset image optimization factor can be configured to make the pixel discarding of the first optimization process and the pixel retention of the second optimization process in an opposite strength state. For example, when the preset image optimization factor is adjusted to strengthen the discarding effect, the pixel retention effect of the second optimization process will be weakened. Among them, the strengthening of the discarding effect corresponds to an increase in the number of pixels selected for discarding (i.e., a larger discarded image region), and the weakening of the pixel retention effect corresponds to a decrease in the number of pixels selected for retention (such as an increase in the sampling interval, etc.). In short, the preset image optimization factor is positively correlated with one of the first quantity and the second quantity and negatively correlated with the other. As the formula provided in the above embodiment, exemplarily, M is in the numerator position in one of the formulas of (X / M)×Y (N = M) in the first optimization process and in the denominator position in the other formula of M×Y (P = M) in the first optimization process, so that the results of the two formulas can tend to change in the opposite direction; when M increases, (X / M)×Y decreases, corresponding to the first optimization process discarding fewer pixels, and correspondingly M×Y increases, corresponding to the second optimization process sampling more pixel numbers. Of course, the above method of setting M in the denominator and numerator of the two formulas to play the opposite strength role of pixel discarding in the first optimization process and pixel retention in the second optimization process is only an example and does not limit its implementation method. In other embodiments, there can also be other solutions, such as constructing the above two formulas through M and I - M (I is a preset value) to control the pixel discarding effect and the pixel retention effect.
[0087] Step S303: Perform stereo matching based on the first optimized image and calculate a first depth map; and, perform stereo matching based on the second optimized image and calculate a second depth image.
[0088] It can be understood that the optimized image pair obtained through the above multi-channel optimization process can be used as the input for various types of stereo matching algorithms. Therefore, the depth measurement method in the embodiments of the present application does not limit the stereo matching algorithms applied, such as various types including region matching, feature matching, and phase matching.
[0089] Step S304: Synthesize the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image.
[0090] In some embodiments, since the first depth image and the second depth image are respectively generated from the first and second optimized image pairs with some original image pixels discarded, the first depth image and / or the second depth image can be restored to the resolution of the original image through an interpolation method, and then the synthesis is performed to obtain a target depth image with the same resolution as the original image.
[0091] Specifically, the interpolation is to infer the approximate depth information at adjacent positions based on the obtained depth information in the first depth image and / or the second depth image, so as to obtain a complete depth image and reduce the pixels with missing depth information. In possible examples, the interpolation method includes but is not limited to the nearest neighbor method, bilinear interpolation method, trilinear interpolation method, etc.
[0092] In possible examples, interpolation can be performed on the second depth image obtained by equally spaced sampling and stereo matching in the above embodiments. For example, for the first original image and the second original image with a size of 400*300, one pixel column is sampled every 1 / 4 width along the horizontal direction of the image, then a second optimized image pair of 4*300 is obtained. After stereo matching and depth calculation, the second depth image is obtained, and then the second depth image is interpolated back to the image resolution of 400*300 between the 4 columns of the second depth image, and then the synthesis is performed.
[0093] In some embodiments, for the synthesis of the target depth image, for each object point, the pixel at the corresponding position with better quality in the first depth image or the second depth image can be selected for filling. For non-corresponding pixels, they can be filled with the pixels at the corresponding positions in one of the restored first depth image and the second depth image. For example, for object point A, the depth information of pixel point A3 corresponding to it in the first depth image is a, and the depth information of pixel point A4 corresponding to it in the second depth image is b. After quality detection, if the quality of A3 is better than that of A4, the value of the corresponding pixel point in the target depth image is filled with a.
[0094] Or, in some other embodiments, the weighted sum of the depth information of the pixel points corresponding to the same object point in the first depth image and the second depth image can also be calculated.
[0095] In possible examples, the detection methods for the quality include but are not limited to at least one of the following:
[0096] 1. Defective pixel detection
[0097] A defective pixel is a pixel point without depth information, and such pixel points can be excluded.
[0098] 2. Outlier detection
[0099] For outlier pixel points whose depth information deviates from a preset threshold (such as 3%), they are untrustworthy abnormal points and can be excluded.
[0100] 3. Noise detection
[0101] Pixel points determined to be image noise can be excluded. For example, fixed pattern noise and temporal noise, etc.
[0102] 4. Image uniformity detection
[0103] For example, brightness uniformity detection and chromaticity uniformity detection, etc., also indicate the change situation of the depth information of each pixel point in the depth image.
[0104] For intuitive illustration Figure 3 of the implementation of the depth measurement method in
[0105] As Figure 5A shown, a flow schematic diagram of the depth measurement method in an application example of the present disclosure is presented.
[0106] In this embodiment, a preset image optimization factor is set as M, and the original image resolutions of the first imaging unit and the second imaging unit are X×Y.
[0107] The first imaging unit captures a first original image, and the second imaging unit captures a second original image. The first original image and the second original image are respectively input into two-way optimization processing, stereo matching, and depth calculation, that is, the first optimization processing 501A and the corresponding stereo matching and depth calculation 502A, as well as the second optimization processing 501B and the corresponding stereo matching and depth calculation 502B. After two-way optimization processing and depth calculation, a first depth image and a second depth image are respectively obtained.
[0108] In 503, based on performing depth measurement quality detection on the first depth image and the second depth image, the first depth image and the second depth image are merged according to the quality detection result to obtain a target depth image.
[0109] Can be combined with Figure 5ASee also Figure 5B and Figure 5C , respectively showing the flow charts of two-way optimization processing and stereo matching and depth calculation.
[0110] exist Figure 5B An application example of the corresponding first optimization process 501A and the corresponding stereo matching and depth calculation 502A is shown in FIG.
[0111] Based on the preset image improvement factor M, the first optimization process is performed on the first original image and the second original image respectively. Exemplarily, Figure 5B The first optimization processing specifically includes: for the first original image and the second original image with an image resolution of X×Y, respectively, centered in the horizontal direction of the image and discarding 1 / M part of the image area, that is, the (X / M)×Y area, to obtain two cropped images with an image resolution of (X×(M-1) / M)×Y, that is, the first optimized image pair.
[0112] Then, by stereo matching and depth information calculation of the first optimized image pair, a first depth image with an image resolution of (X×(M-1) / M)×Y is obtained.
[0113] exist Figure 5C An application example of the corresponding second optimization process 501B and the corresponding stereo matching and depth calculation 502B is shown in FIG.
[0114] Based on the preset image improvement factor M, the second optimization process is performed on the first original image and the second original image respectively. Exemplarily, Figure 5C The second optimization processing specifically includes: for the first original image and the second original image with an image resolution of X×Y, 1 / M sampling and discarding processing is performed in the horizontal direction of the image; that is, sampling is started from the first leftmost column of pixels, sampling once every M pixels and discarding the image area between the sampled columns of pixels to obtain two sampled images with an image resolution of (X / M)×Y, that is, the second optimized image pair.
[0115] Then, stereo matching is performed on the first optimized image pair and the depth is calculated to obtain a depth image with an image resolution of (X / M)×Y. Figure 5B Slightly different is that in Figure 5C In the process, interpolation is also performed on this (X / M)×Y local depth image to reconstruct the depth information corresponding to the discarded image area, thereby obtaining a second depth image with an image resolution of X×Y.
[0116] For example, based on Figure 5BThe first X×Y depth image with an image resolution of (X×(M - 1) / M)×Y and the second depth image with an image resolution of X×Y are obtained, depth information quality detection is performed, and the two depth images are merged according to the quality detection result to obtain a target depth image for output. Among them, the depth information missing in the image area corresponding to the centrally cropped image area of the first depth image in the target depth image can be filled by the second depth image; the depth information in the two side areas outside the centrally cropped image area in the target image depth image can be filled according to the more reliable depth information among the "original" depth information corresponding to the first depth image and the interpolated "estimated" depth information corresponding to the second depth image.
[0117] As Figure 6A and Figure 6B shown, Figure 6A A schematic diagram showing the depth image obtained when the depth measurement method of the present disclosure is not executed, where the black area indicates missing depth information, corresponding to the failure of stereo matching in some relatively close areas such as the ground. Again, as Figure 6B shown, a schematic diagram showing the depth image obtained when the depth measurement method of the present disclosure is executed can be found that the black area is effectively reduced, effectively improving the quality of the depth image.
[0118] It should be specifically noted that although the "height" of the binocular stereo vision system is described relative to the ground in the above examples, in fact, the ground is only a bearing surface provided in an application scenario (such as indoors) of a carrier (such as a sweeping robot) of the binocular stereo vision system, and the carrier can move on the bearing surface.
[0119] In other scenarios, the bearing surface can change according to different scenarios. For example, for a robot moving on a wall, the bearing surface is the wall. There is also the problem that the closer the binocular stereo vision system is to the wall, the greater the parallax. This problem can be solved by the depth measurement method provided in the embodiments of the present disclosure. Therefore, it can be understood that the ground is not a limitation.
[0120] In addition, the example of less than 20 cm given in the foregoing examples is only an example. In fact, the value of this height is relative. Specifically, relative to the detection distance of the binocular stereo vision system being 10 meters, a height within 20 cm can be recognized as close to the bearing plane; while relative to the detection distance of 200 meters, the height of the binocular stereo vision system "close" to the bearing plane will be raised to a certain extent, such as 1 meter, etc. Therefore, the distance between the binocular stereo vision system and the bearing surface is positively correlated with the detection distance of the binocular stereo vision system.
[0121] As Figure 7As shown, it is a schematic diagram of the modules of the depth measurement device in an embodiment of the present disclosure. It should be noted that the principle of the depth measurement device can refer to the depth measurement method in the previous embodiment, so the same technical content will not be repeated here.
[0122] The depth measurement device 700 is applied to a binocular stereo vision system, and it includes:
[0123] An image acquisition module 701, configured to acquire the original image pair captured by the binocular stereo vision system;
[0124] A multi-channel optimization module 702, configured to perform a first optimization process and a second optimization process on the basis of the original image pair respectively to obtain a first optimized image pair and a second optimized image pair; wherein, the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image; the processed image is each original image in the original image pair;
[0125] A multi-channel depth calculation module 703, configured to perform stereo matching on the basis of the first optimized image pair and calculate to obtain a first depth map; and perform stereo matching on the basis of the second optimized image pair and calculate to obtain a second depth image;
[0126] A synthesis module 704, configured to synthesize the first depth image and the second depth image on the basis of the depth measurement quality result to obtain a target depth image.
[0127] In some embodiments, the first optimization process and the second optimization process are respectively image processing methods of selectively discarding and selectively retaining some pixels of the processed image in the horizontal direction of the image; the first quantity of the part of the pixels selectively discarded by the first optimization process is determined by a first image optimization factor; the second quantity of the part of the pixels selectively retained by the second optimization process is determined by a second image optimization factor; the first image optimization factor and the second image optimization factor are the same or different.
[0128] In some embodiments, the first image optimization factor and the second image optimization factor are the same preset image optimization factor.
[0129] In some embodiments, the preset image optimization factor is used to control the pixel discarding effect of the first optimization process and the pixel retaining effect of the second optimization process in a strong and weak opposite manner.
[0130] In some embodiments, the preset image optimization factor is positively correlated with one of the first quantity and the second quantity, and negatively correlated with the other.
[0131] In some embodiments, before synthesizing the first depth image and the second depth image, it further includes: restoring the first depth image and / or the second depth image to the original image resolution by interpolation.
[0132] In some embodiments, the first optimization process includes: cropping and discarding an image region from the processed image to obtain an optimized image for stereo matching; wherein, the vertical image resolution between the cropped image region and the processed image is the same, and the ratio of the horizontal image resolution of the cropped image region to the processed image is 1:N, where N is an integer greater than 1; and / or, the second optimization process includes: equally sampling a pixel column from the processed image along the horizontal direction of the image and discarding the image regions between the pixel columns to obtain an optimized image for stereo matching; wherein, the pixel column has the same vertical image resolution as the processed image, and the ratio of the interval to the width of the processed image in the horizontal direction of the image is 1:P, where P is an integer greater than 1.
[0133] In some embodiments, both N and P are the image optimization factor M.
[0134] In some embodiments, the cropping and discarding an image region from the processed image includes: centrally cropping and discarding the image region from the processed image.
[0135] In some embodiments, before synthesizing the first depth image and the second depth image, it further includes: performing image interpolation based on the second depth image to restore it to the original image resolution.
[0136] In some embodiments, the synthesis module 704 is configured to, for each pixel point of the object point in the target depth image, select the pixel point at the corresponding position with better quality in the first depth image or the second depth image for filling.
[0137] In some embodiments, the binocular stereo vision system is arranged such that its optical axis direction is not perpendicular to a bearing surface; wherein, the carrier of the binocular stereo vision system is movably disposed on the bearing surface.
[0138] In some embodiments, the binocular stereo vision system is arranged close to the bearing surface, and the distance between the binocular stereo vision system and the bearing surface is positively correlated with the detection distance of the binocular stereo vision system.
[0139] It should be specifically noted that in Figure 7Each functional module in the embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a program instruction product. The program instruction product includes one or more program instructions. When the program instructions are loaded and executed on a computer, the processes or functions according to the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The program instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.
[0140] And, Figure 7 The devices disclosed in the embodiments can be implemented by other module division methods. The device embodiments shown above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or modules can be combined or can be dynamically moved to another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical or other forms.
[0141] In addition, Figure 7 Each functional module and sub-module in the embodiments can be dynamically located in a processing component, or each module can exist physically alone, or two or more modules can be dynamically located in a component. The above-mentioned dynamic component can be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned dynamic component is implemented in the form of a software functional module and executed as an independent product for sale or use, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a disk, an optical disc, etc.
[0142] As Figure 8 shown, a schematic structural diagram of a binocular stereo vision system shown in an embodiment of the present disclosure is presented. The binocular stereo vision system 800 can be packaged as a device or a component for use.
[0143] The binocular stereo vision system 800 includes:
[0144] One or a pair of imaging units for collecting a pair of original images with parallax. Figure 8 Exemplarily shown in the figure as a pair of first imaging unit 801 and second imaging unit 802.
[0145] A processing unit 803, communicatively connected to the first imaging unit 801 and the second imaging unit 802, for running program instructions to perform steps such as Figure 3 those described in the depth measurement method.
[0146] Optionally, if it is an active ranging binocular stereo vision system, it may further include an infrared projector.
[0147] In some embodiments, the processing unit 802 may be implemented by a processor and a memory. The memory stores program instructions, and the processor is configured to run the program instructions. Exemplarily, the memory may be implemented based on volatile and / or non-volatile memory, and the processor may be implemented by a central processing unit, a SoC, or other processors.
[0148] In an embodiment of the present disclosure, a motion device may also be provided. The motion device includes: a motion device body movably disposed on a bearing surface; for example Figure 7 The binocular stereo vision system is disposed on the motion device body and is configured such that the optical axis direction thereof is not perpendicular to the bearing surface. In a possible example, the motion device may be implemented as, for example, a mobile robot, such as various transport robots, service robots, etc.; or it may include a vehicle, such as a vehicle with an autonomous driving function, etc.
[0149] In an embodiment of the present disclosure, a computer-readable storage medium may also be provided, storing program instructions, and the program instructions, when run, execute the method steps in the previous Figure 3 embodiment.
[0150] The method steps in the above embodiments are implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or are implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium and downloaded through a network, so that the method represented herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA).
[0151] It should be specifically noted that the flowcharts representing the processes or methods in the above embodiments of the present disclosure can be understood as representing modules, segments, or portions of code including one or more executable instructions for implementing specific logical functions or processes. And the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed.
[0152] For example, Figure 3 the order of the steps in the embodiments may be changed in a specific scenario, not limited to the above representation.
[0153] Such as Figure 9As shown, it is a schematic structural diagram of a computer device in an embodiment of the present disclosure.
[0154] In some embodiments, the computer device is used to load program instructions for implementing the depth measurement method. The computer device can be specifically implemented as, for example, a server, a desktop computer, a laptop computer, a mobile terminal, etc., and may be used by implementers who store and / or run these program instructions for commercial purposes such as development and testing.
[0155] Figure 9 The shown computer device 900 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0156] As Figure 9 shown, the computer device 900 is presented in the form of a general-purpose computing device. The components of the computer device 900 may include, but are not limited to: at least one of the above-mentioned processing units 910, at least one of the above-mentioned storage units 920, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).
[0157] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 910, so that the computer device is used to implement the method steps described in the above embodiments of the present disclosure (for example Figure 3 embodiments).
[0158] In some embodiments, the storage unit 920 may include a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202, and may further include a read-only storage unit (ROM) 9203.
[0159] In some embodiments, the storage unit 920 may further include a program / utility 9204 having a set (at least one) of program modules 9205. Such program modules 9205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0160] In some embodiments, the bus 930 may include a data bus, an address bus, and a control bus.
[0161] In some embodiments, the computer device 900 may also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and such communication may be performed through the input / output (I / O) interface 950. Optionally, the computer device 900 further includes a display unit 940, which is connected to the input / output (I / O) interface 950 for display. Moreover, the computer device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 9100. As shown in the figure, the network adapter 9100 communicates with other modules of the computer device 900 through the bus 930. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computer device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0162] In summary, the embodiments of the present disclosure provide a depth measurement method, a binocular stereo vision system, a device, and a storage medium. First optimization processing and second optimization processing are respectively performed on original image pairs to obtain a first optimized image pair and a second optimized image pair; wherein, the first optimization processing and the second optimization processing are different image processing methods for reducing the pixels of the processed images in the horizontal direction of the images; stereo matching is performed based on the first optimized image pair and a first depth map is calculated; and, stereo matching is performed based on the second optimized image pair and a second depth image is calculated; the first depth image and the second depth image are synthesized based on the depth measurement quality result to obtain a target depth image. By performing multi-path optimization processing on the original image pairs captured by the binocular stereo system to reduce the reduced parallax and then performing stereo matching to obtain depth images and synthesize them, the influence of parallax is reduced to improve the success rate of stereo matching and the detection ability.
[0163] The above embodiments are only illustrative of the principles and effects of the present disclosure, and are not used to limit the present disclosure. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present disclosure. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present disclosure should still be covered by the claims of the present disclosure.
Claims
1. A depth measurement method, characterized in that, Applied to a binocular stereo vision system: The depth measurement method includes: Obtaining an original image pair captured by the binocular stereo vision system; Performing a first optimization process and a second optimization process on the basis of the original image pair respectively to obtain a first optimized image pair and a second optimized image pair; wherein, the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image; the processed image is each original image in the original image pair; Performing stereo matching on the basis of the first optimized image pair and calculating to obtain a first depth map; and performing stereo matching on the basis of the second optimized image pair and calculating to obtain a second depth image; Synthesizing the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image; The first optimization process and the second optimization process are performed based on at least one adjustable image optimization factor, the image optimization factor is a preset image optimization factor, and the preset image optimization factor is used to control the pixel discarding effect of the first optimization process and the pixel retaining effect of the second optimization process in the opposite strength, so as to adjust the disparity between the paired corresponding pixels in each of the first optimized image pair and the second optimized image pair.
2. The depth measurement method according to claim 1, characterized in that, The first optimization process and the second optimization process are respectively image processing methods of selecting to discard and selecting to retain some pixels of the processed image in the horizontal direction of the image; the first quantity of the part of the pixels selected to be discarded by the first optimization process is determined by a first image optimization factor; the second quantity of the part of the pixels selected to be retained by the second optimization process is determined by a second image optimization factor; the first image optimization factor and the second image optimization factor are the same or different.
3. The depth measurement method according to claim 2, characterized in that, The first image optimization factor and the second image optimization factor are the same preset image optimization factor.
4. The depth measurement method according to claim 3, wherein The preset image optimization factor is positively correlated with one of the first quantity and the second quantity and negatively correlated with the other.
5. The depth measurement method according to claim 1, wherein Before synthesizing the first depth image and the second depth image, it further includes: Restoring the first depth image and / or the second depth image to the original image resolution by interpolation.
6. The depth measurement method according to claim 1, characterized in that, The first optimization process includes: cropping and discarding an image area from the processed image to obtain an optimized image for stereo matching; wherein, the vertical image resolution between the cropped image area and the processed image is the same, and the ratio of the horizontal image resolution of the cropped image area to the processed image is 1:N, and N is an integer greater than 1; And / or, the second optimization process includes: equally spaced sampling a pixel column from the processed image along the horizontal direction of the image and discarding the image area between each pixel column to obtain an optimized image for stereo matching; wherein, the vertical image resolution of the pixel column is the same as that of the processed image, and the ratio of the interval to the width of the processed image in the horizontal direction of the image is 1:P, and P is an integer greater than 1.
7. The depth measurement method according to claim 6, characterized in that, Both N and P are the image optimization factor M.
8. The depth measurement method according to claim 6, wherein The cropping and discarding an image area from the processed image includes: centrally cropping and discarding the image area of the processed image.
9. The depth measurement method according to claim 6, wherein Before synthesizing the first depth image and the second depth image, it further includes: Perform image interpolation based on the second depth image to restore it to the original image resolution.
10. The depth measurement method according to claim 1, wherein Synthesize the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image, including: For each pixel point of the target depth image corresponding to an object point, select the pixel point at the corresponding position with better quality in the first depth image or the second depth image for filling.
11. The depth measurement method according to claim 1, characterized in that, The binocular stereo vision system is set such that the optical axis direction is not perpendicular to a bearing surface; wherein, the carrier of the binocular stereo vision system is movably disposed on the bearing surface.
12. The depth measurement method according to claim 11, characterized in that, The binocular stereo vision system is set to be close to the bearing surface, and the distance between the binocular stereo vision system and the bearing surface is positively correlated with the detection distance of the binocular stereo vision system.
13. A depth measurement device, characterized in that, Applied to a binocular stereo vision system: The depth measurement device includes: An image acquisition module, configured to acquire an original image pair captured by the binocular stereo vision system. A multi-channel optimization module, configured to perform a first optimization process and a second optimization process on the original image pair respectively to obtain a first optimized image pair and a second optimized image pair; wherein, the first optimization process and the second optimization process are different image processing methods for reducing the pixels of the processed image in the horizontal direction of the image; the processed image is each original image in the original image pair. A multi-channel depth calculation module, configured to perform stereo matching based on the first optimized image pair and calculate a first depth map; and perform stereo matching based on the second optimized image pair and calculate a second depth image. A synthesis module, configured to synthesize the first depth image and the second depth image based on the depth measurement quality result to obtain a target depth image. The first optimization process and the second optimization process are performed based on at least one adjustable image optimization factor, the image optimization factor is a preset image optimization factor, and the preset image optimization factor is used to control the pixel discarding effect of the first optimization process and the pixel retaining effect of the second optimization process in opposite strengths, so as to adjust the disparity between the paired corresponding pixel points in each of the first optimized image pair and the second optimized image pair.
14. A binocular stereo vision system, characterized in that, Including: One or a pair of camera units, configured to collect an original image pair with disparity. A processing unit, communicatively connected to each camera unit, configured to run program instructions to perform the depth measurement method according to any one of claims 1 to 12.
15. A sports device, characterized in that, Including: A moving device body, movably disposed on a bearing surface. The binocular stereo vision system according to claim 14, disposed on the moving device body, and set such that the optical axis direction is not perpendicular to the bearing surface.
16. The exercise device according to claim 15, wherein, The optical axis direction of the binocular stereo vision system is set to correspond to the front, rear or side of the moving device body.
17. The exercise device according to claim 15, wherein, The bearing surface is the ground, and the moving device includes a mobile robot.
18. A computer device, characterized in that, Including: A communicator, a memory and a processor. The communicator is used for external communication; the memory stores program instructions; the processor is used to run the program instructions to perform the depth measurement method according to any one of claims 1 to 12.
19. A computer-readable storage medium, characterized in that, A program instruction is stored, and the program instruction is run to execute the depth measurement method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing method and device
CN108496201A