Stereoscopic depth recognition method and device, control device, and readable storage medium

By disparity processing and optimization of the binocular scene map taken by the robot, the body depth distribution map is created, which solves the problem of insufficient recognition universality among different robots in the prior art, and achieves high-precision and high-efficiency three-dimensional depth recognition.

CN114821330BActive Publication Date: 2025-06-27UBTECH ROBOTICS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210517451.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-06-27
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

When used between robots in different styles and scenarios, the existing three-dimensional depth recognition scheme is insufficient in versatility and cannot be effectively compatible, resulting in low recognition accuracy and efficiency.

Method used

By obtaining the left and right eye scene maps taken by the target robot, the initial parallax map is determined, the parallax range classification and optimization are performed, the pre-stored parallax regression model is called for processing, and finally the parallax depth conversion is performed based on the camera focal length and baseline distance to create a body depth distribution map.

Benefits of technology

It realizes high-precision and efficient three-dimensional depth recognition on robots of different styles and scenarios, reducing the calculation amount and improving recognition accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821330B_ABST
    Figure CN114821330B_ABST
Patent Text Reader

Abstract

The present application provides a stereo depth recognition method, device, control equipment and readable storage medium, relating to the field of computer vision technology. After determining an initial disparity map between a left-eye scene map and a right-eye scene map captured by a target robot for a to-be-measured scene, the present application can classify each disparity value in the initial disparity map into a disparity range, update the obtained disparity range distribution result according to a preset expected normal distribution, then call a pre-stored disparity regression model according to the updated disparity distribution optimization result to perform disparity regression processing on the initial disparity map, and finally perform disparity depth conversion processing on the regressed target scene disparity map to obtain a stereo depth distribution map between the target robot and the to-be-measured scene, so as to ensure that the entire stereo depth recognition solution can be effectively compatible with robots of different models / scenes through the disparity range distribution update operation to achieve a high-precision and high-efficiency stereo depth recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and more particularly, to a stereo depth recognition method and apparatus, a control device, and a readable storage medium. Background Art

[0002] With the continuous development of science and technology, computer vision technology has received extensive attention from all walks of life due to its great research value and application value. Among them, the robot moving obstacle avoidance function is an important application direction of computer vision technology. It is often necessary to use computer vision technology to process the binocular scene view captured by the robot to determine the true stereo depth condition between the captured scene and the robot, so as to facilitate the robot to avoid obstacles and move.

[0003] It should be noted that the camera baseline distance between the left eye camera and the right eye camera of different models of robots often varies, and the sizes of different models of robots (e.g., robot height and / or robot width) are not exactly the same, resulting in obvious differences in the binocular scene views captured by different models of robots for the same scene. At the same time, the obstacle avoidance scenarios that different models of robots are good at are not exactly the same, which will inevitably lead to poor generality of the stereo depth recognition schemes matched by different models of robots and cannot be well compatible with robots of different models / scenarios for use. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a stereo depth recognition method and apparatus, a control device, and a readable storage medium, which can be effectively compatible with robots of different models / scenarios to achieve high-precision and high-efficiency stereo depth recognition effects.

[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, this application provides a stereo depth recognition method, and the method includes:

[0007] Obtain a left-eye scene map and a right-eye scene map captured by a target robot for a to-be-measured scene;

[0008] Determine an initial disparity map of the right-eye scene map relative to the left-eye scene map;

[0009] Classify the disparity values in the initial disparity map into disparity range categories to obtain a corresponding disparity range distribution result;

[0010] Update the disparity range distribution result of the initial disparity map according to a preset expected normal distribution to obtain a corresponding disparity distribution optimization result;

[0011] Call the pre - stored disparity regression model according to the optimized result of the disparity distribution, perform disparity regression processing on the initial disparity map, and obtain a target scene disparity map matching the to - be - measured scene;

[0012] According to the camera focal length and camera baseline distance of the target robot, perform disparity - depth conversion processing on the target scene disparity map to obtain a three - dimensional depth distribution map between the target robot and the to - be - measured scene.

[0013] In an alternative embodiment, the step of determining the initial disparity map of the right - eye scene map relative to the left - eye scene map includes:

[0014] Extract features from the right - eye scene map and the left - eye scene map respectively to obtain corresponding right - eye scene feature maps and left - eye scene feature maps;

[0015] Perform pixel - feature matching between the left - eye scene feature map and the right - eye scene feature map to obtain the disparity values between each pixel point in the left - eye scene feature map and the feature - matching pixel points in the right - eye scene feature map;

[0016] Perform image aggregation on the disparity values corresponding to all pixel points in the left - eye scene feature map to obtain the initial disparity map.

[0017] In an alternative embodiment, the step of classifying the disparity values in the initial disparity map into disparity - range categories to obtain a corresponding disparity - range distribution result includes:

[0018] Determine the disparity - range specification corresponding to the initial disparity map according to the maximum disparity value in the initial disparity map and the preset disparity - range planning step size;

[0019] Call the cross - entropy loss function according to the disparity - range specification to identify the disparity ranges to which all disparity values in the initial disparity map belong;

[0020] According to the occurrence frequencies of different disparity ranges corresponding to the initial disparity map, fit the distribution probability curve of the disparity ranges at the initial disparity map to obtain the disparity - range distribution result.

[0021] In an alternative embodiment, the step of updating the disparity - range distribution result of the initial disparity map according to a preset expected normal distribution to obtain a corresponding optimized disparity - distribution result includes:

[0022] Judge whether the standard deviation value of the normal - distribution curve corresponding to the disparity - range distribution result is greater than the preset standard - deviation threshold;

[0023] If it is determined that the standard deviation value is greater than the preset standard deviation threshold, calculate the mathematical expectation difference and the standard deviation multiple between the expected normal distribution and the normal distribution curve, and perform a curve shift on the normal distribution curve corresponding to the parallax range distribution result according to the mathematical expectation difference and the standard deviation multiple to obtain the optimized parallax distribution result;

[0024] If it is determined that the standard deviation value is less than or equal to the preset standard deviation threshold, directly use the parallax range distribution result as the optimized parallax distribution result.

[0025] In an alternative embodiment, the step of calling a pre-stored parallax regression model according to the optimized parallax distribution result to perform parallax regression processing on the initial parallax map to obtain a target scene parallax map matching the to-be-detected scene includes:

[0026] Based on the initial parallax map, call the parallax regression model to estimate all possible parallax values and the occurrence probabilities of all possible parallax values for each pixel point in the target scene parallax map that satisfy the optimized parallax distribution result;

[0027] For each pixel point in the target scene parallax map, perform a weighted summation operation on all possible parallax values corresponding to the pixel point and the occurrence probabilities of all possible parallax values to obtain the target scene parallax value at the pixel point in the target scene parallax map.

[0028] In an alternative embodiment, the step of performing parallax depth conversion processing on the target scene parallax map according to the camera focal length and the camera baseline distance of the target robot to obtain a stereo depth distribution map between the target robot and the to-be-detected scene includes:

[0029] For each pixel point in the target scene parallax map, substitute the camera focal length, the camera baseline distance, and the target scene parallax value corresponding to the pixel point into a pre-stored parallax depth correlation relationship to calculate the target scene depth value corresponding to the pixel point at the stereo depth distribution map;

[0030] Perform image aggregation on the target scene depth values corresponding to all pixel points in the target scene parallax map to obtain the stereo depth distribution map.

[0031] In an alternative embodiment, the parallax depth correlation relationship is expressed by the following formula:

[0032]

[0033] Wherein, D is used to represent the parallax value between the left-eye camera and the right-eye camera of the target robot that are on the same baseline for the object to be photographed, b is used to represent the camera baseline distance between the left-eye camera and the right-eye camera of the target robot, f is used to represent the camera focal length of the target robot, and Z is used to represent the stereo depth value between the object to be photographed and the target robot.

[0034] In a second aspect, the present application provides a stereo depth recognition device, and the device includes:

[0035] A scene view acquisition module, configured to acquire a left-eye scene map and a right-eye scene map photographed by the target robot for the scene to be measured;

[0036] An initial parallax determination module, configured to determine an initial parallax map of the right-eye scene map relative to the left-eye scene map;

[0037] A parallax distribution classification module, configured to classify the parallax values in the initial parallax map into parallax range classifications to obtain a corresponding parallax range distribution result;

[0038] A parallax distribution optimization module, configured to update the parallax range distribution result of the initial parallax map according to a preset expected normal distribution to obtain a corresponding parallax distribution optimization result;

[0039] A scene parallax regression module, configured to call a pre-stored parallax regression model according to the parallax distribution optimization result to perform parallax regression processing on the initial parallax map to obtain a target scene parallax map matching the scene to be measured;

[0040] A stereo depth calculation module, configured to perform parallax depth conversion processing on the target scene parallax map according to the camera focal length and the camera baseline distance of the target robot to obtain a stereo depth distribution map between the target robot and the scene to be measured.

[0041] In a third aspect, the present application provides a control device, including a processor and a memory, the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the stereo depth recognition method described in any one of the foregoing embodiments.

[0042] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the stereo depth recognition method described in any one of the foregoing embodiments is implemented.

[0043] In this case, the beneficial effects of the embodiments of the present application include the following:

[0044] After determining the initial disparity map between the left-eye scene map and the right-eye scene map captured by the target robot for the to-be-measured scene, the present application can classify the disparity values in the initial disparity map into disparity ranges to obtain a disparity range distribution result, update the disparity range distribution result according to a preset expected normal distribution to obtain an optimized disparity distribution result, and then call a pre-stored disparity regression model according to the optimized disparity distribution result to perform disparity regression processing on the initial disparity map to obtain a target scene disparity map matching the to-be-measured scene. Then, according to the camera focal length and the camera baseline distance of the target robot, perform disparity depth conversion processing on the target scene disparity map to obtain a three-dimensional depth distribution map between the target robot and the to-be-measured scene, so as to migrate the image data to be recognized in different situations to the same expected data domain for three-dimensional depth recognition through the disparity range distribution update operation, reduce the data calculation amount of the entire depth recognition solution, improve the universality and recognition accuracy of the depth recognition solution, and ensure that the three-dimensional depth recognition solution provided by the present application can be effectively compatible with robots of different models / scenes to achieve a high-precision and high-efficiency three-dimensional depth recognition effect.

[0045] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0047] Figure 1 Schematic diagram of the composition of the control device provided for the embodiments of the present application;

[0048] Figure 2 Schematic diagram of the flow of the three-dimensional depth recognition method provided for the embodiments of the present application;

[0049] Figure 3 For Figure 2 Schematic diagram of the sub-steps included in step S220 in

[0050] Figure 4 For Figure 2 Schematic diagram of the sub-steps included in step S230 in

[0051] Figure 5 For Figure 2 Schematic diagram of the sub-steps included in step S240 in

[0052] Figure 6 It is Figure 2 a schematic flow diagram of the sub-steps included in step S250 in

[0053] Figure 7 It is Figure 2 a schematic flow diagram of the sub-steps included in step S260 in

[0054] Figure 8 a diagram showing the effect of the parallax depth correlation relationship provided by the embodiment of the present application;

[0055] Figure 9 a schematic diagram of the composition of the stereo depth recognition device provided by the embodiment of the present application.

[0056] Icons: 10 - control device; 11 - memory; 12 - processor; 13 - communication unit; 100 - stereo depth recognition device; 110 - scene view acquisition module; 120 - initial parallax determination module; 130 - parallax distribution classification module; 140 - parallax distribution optimization module; 150 - scene parallax regression module; 160 - stereo depth calculation module. Detailed implementation manners

[0057] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0058] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0059] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0060] In the description of the present application, it should be understood that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood in specific circumstances.

[0061] Through painstaking research, the applicant found that existing stereo depth recognition solutions need to combine the binocular camera configuration of the corresponding robot (including the position distribution of the left and right eye cameras respectively, and the baseline distance between the left and right eye cameras) and the binocular scene view captured by the robot in a specific training scenario on the basis of a conventional convolutional neural network model for model training, so as to realize the stereo depth recognition function of the robot with a specific binocular camera configuration for a specific scene through the trained stereo depth recognition model. However, it should be noted that when the constructed stereo depth recognition model is applied to other binocular camera configurations and / or other scenarios, it often fails to match the attributes of the model training data (for example, the camera baseline distance and training scene corresponding to the stereo depth recognition model), resulting in low accuracy of the recognized stereo depth data. At the same time, due to the model characteristics of the convolutional neural network model, the trained stereo depth recognition model has the characteristics of high computational complexity, and a large amount of computational power must be consumed to realize the corresponding function, and the overall recognition efficiency is poor.

[0062] In this case, to provide a stereo depth recognition solution that can be effectively compatible with robots of different models / scenarios to achieve high-precision and high-efficiency stereo depth recognition functions and ensure the universality of the depth recognition solution, the embodiments of the present application implement the foregoing functions by providing a stereo depth recognition method and device, a control device, and a readable storage medium.

[0063] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0064] Please refer to Figure 1 , Figure 1It is a schematic diagram of the composition of the control device 10 provided by the embodiments of the present application. In this embodiment, the control device 10 can be communicatively connected to at least one robot and control the at least one robot to move to a certain or several scenarios respectively for image acquisition, so that the control device 10 can determine the stereo depth information of each obstacle in the corresponding scenario from the stereo scene views collected by each robot for each robot, and control the corresponding robot to perform obstacle avoidance movement according to the determined stereo depth information. Among them, the binocular camera configuration conditions of different robots can be the same or different. For example, the camera baseline distances of different robots are not the same or partially the same; the scenarios where different robots are located can be the same or different. During this process, the control device 10 can provide a general stereo depth recognition solution with strong universality for robots of different models / scenarios, and achieve a high-precision and high-efficiency stereo depth recognition effect through this stereo depth recognition solution.

[0065] In addition, the control device 10 can also be separately deployed on a single robot, so that when the robot replaces the binocular camera configuration and / or the scenario where it is located, it can provide a general stereo depth recognition solution with strong universality for the robot to maintain a high-precision and high-efficiency stereo depth recognition effect for a long time.

[0066] In this embodiment, the control device 10 may include a memory 11, a processor 12, a communication unit 13, and a stereo depth recognition device 100. Among them, the memory 11, the processor 12, and the communication unit 13 are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these elements such as the memory 11, the processor 12, and the communication unit 13 can be electrically connected to each other through one or more communication buses or signal lines.

[0067] In this embodiment, the memory 11 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory 11 is used to store a computer program, and after receiving an execution instruction, the processor 12 can execute the computer program accordingly.

[0068] In this embodiment, the processor 12 may be an integrated circuit chip with signal processing capability. The processor 12 may be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or at least one of other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc., which may implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application.

[0069] In this embodiment, the communication unit 13 is used to establish a communication connection between the control device 10 and each robot through a network, and to send and receive data through the network, wherein the network includes a wired communication network and a wireless communication network. For example, the control device 10 can control each robot to perform an image acquisition operation and control each robot to perform a movement operation through the communication unit 13.

[0070] In this embodiment, the stereo depth recognition device 100 includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or fixed in the operating system of the control device 10. The processor 12 can be used to execute the executable modules stored in the memory 11, such as the software function modules and computer programs included in the stereo depth recognition device 100. The control device 10 can effectively be compatible with robots of different styles / scenes through the stereo depth recognition device 100 to achieve high-precision and high-efficiency stereo depth recognition effects and ensure depth recognition accuracy.

[0071] Understandably, Figure 1 The block diagram shown is only a schematic diagram of a composition of the control device 10. The control device 10 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0072] In the present application, in order to ensure that the control device 10 can provide a stereo depth recognition solution that is compatible with robots of different styles / scenes and realizes a high-precision and high-efficiency stereo depth recognition function, the present embodiment of the application achieves the above-mentioned purpose by providing a stereo depth recognition method. The stereo depth recognition method provided by the present application is described in detail below.

[0073] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the stereo depth recognition method provided by an embodiment of the present application. In the embodiment of the present application, the stereo depth recognition method may include steps S210 to S260.

[0074] Step S210: Obtain a left-eye scene image and a right-eye scene image captured by the target robot for the to-be-measured scene.

[0075] In this embodiment, the target robot is the currently selected robot that needs to confirm the distance (i.e., stereo depth information) between it and each obstacle in the current scene, and the to-be-measured scene is the scene where the target robot is currently located and requires stereo depth recognition processing. The control device 10 can obtain, through the communication unit 13, the left-eye scene image captured by the left-eye camera of the target robot for the to-be-measured scene, and the right-eye scene image captured by the right-eye camera of the target robot for the to-be-measured scene.

[0076] Step S220: Determine an initial disparity map of the right-eye scene image relative to the left-eye scene image.

[0077] In this embodiment, after the control device 10 obtains the left-eye scene image and the right-eye scene image captured by the target robot for the to-be-measured scene, it can perform data processing on the left-eye scene image and the right-eye scene image by invoking a pre-stored disparity construction model to construct an initial disparity map of the right-eye scene image relative to the left-eye scene image; or it can perform feature extraction on the left-eye scene image and the right-eye scene image by invoking a pre-stored feature extraction model, and then determine the disparity values at each pixel position in the initial disparity map through feature matching. At this time, the initial disparity map represents the distribution result of the disparity values of the right-eye scene image relative to the left-eye scene image estimated preliminarily.

[0078] Step S230: Classify the disparity values in the initial disparity map into disparity ranges to obtain a corresponding disparity range distribution result.

[0079] In this embodiment, after the control device 10 determines the initial disparity map between the right-eye scene map and the left-eye scene map of the target robot for the to-be-measured scene, it can construct a plurality of continuous disparity ranges related to the initial disparity map according to the maximum disparity value in the initial disparity map. Each disparity range represents a disparity value distribution interval. At this time, the control device 10 can classify the disparity values in the initial disparity map to determine which disparity range each disparity value in the initial disparity map belongs to, and then obtain a disparity range distribution result for characterizing the distribution probability status of each disparity range involved in the initial disparity map according to the occurrence frequency of each disparity range in the initial disparity map.

[0080] In an implementation manner of this embodiment, to improve the disparity range classification efficiency of the control device 10 and reduce the classification calculation amount of the control device 10, each disparity value in the determined initial disparity map can be divided by the same positive integer greater than 1, and the calculated value is used to replace the corresponding disparity value in the initial disparity map, so as to effectively reduce the data operation amount of the control device 10 for performing the disparity range classification operation and improve the disparity range classification efficiency.

[0081] Step S240: Update the disparity range distribution result of the initial disparity map according to a preset expected normal distribution to obtain a corresponding disparity distribution optimization result.

[0082] In this embodiment, the expected normal distribution is used to characterize the lowest expected status of the user for the stereo depth recognition effect, and it describes the mathematical expectation and standard deviation of the disparity range distribution status required to achieve the expected depth recognition effect when represented by a normal distribution curve. The control device 10 updates the disparity range distribution result of the initial disparity map according to the obtained expected normal distribution to migrate the disparity data directly matching the captured binocular scene map of the target robot into the preset expected data domain, ensuring that the finally obtained disparity distribution optimization result can effectively reduce the data operation amount of the control device 10 for performing the disparity regression operation and improve the adaptability of the control device 10 to different camera configurations and / or robot usage scenarios.

[0083] Step S250: Invoke a pre-stored disparity regression model according to the disparity distribution optimization result to perform disparity regression processing on the initial disparity map to obtain a target scene disparity map matching the to-be-measured scene.

[0084] In this embodiment, after the control device 10 determines the optimized parallax distribution result, it can perform parallax regression processing on the initial parallax map based on the optimized parallax distribution result by invoking a parallax regression model, so as to use the parallax value data migrated into the expected data domain to perform a high-precision parallax prediction operation, and obtain a target scene parallax map of the target robot that matches the scene to be measured with high precision.

[0085] In an implementation manner of this embodiment, if a parallax value reduction operation is performed on the initial parallax map before performing the parallax range classification operation (that is, the value obtained by dividing each parallax value in the determined initial parallax map by the same positive integer greater than 1 is used to replace the corresponding original parallax value in the initial parallax map), then after the above parallax regression operation is completed, the same positive integer greater than 1 needs to be used to magnify each target scene parallax value in the obtained target scene parallax map, and the magnified target scene parallax map is used as the target scene parallax map for the control device 10 to perform the parallax depth conversion operation.

[0086] Step S260: Perform parallax depth conversion processing on the target scene parallax map according to the camera focal length and camera baseline distance of the target robot, to obtain a stereo depth distribution map between the target robot and the scene to be measured.

[0087] In this embodiment, after the control device 10 obtains a target scene parallax map of the target robot that matches the scene to be measured with high precision, the control device 10 will perform parallax depth conversion processing on the high-precision target scene parallax map according to the camera baseline distance between the current binocular cameras of the target robot and the camera focal length when the target robot currently captures images of the scene to be measured, to obtain a stereo depth distribution map between the target robot and the scene to be measured that satisfies the parallax depth association relationship. Among them, the parallax depth association relationship is used to represent the conversion relationship between the stereo depth and parallax value between the binocular camera and the photographed object.

[0088] Therefore, this application can perform the above steps S210 to S260, migrate the to-be-recognized image data in different situations (including different camera configurations and / or different scenes) to the same expected data domain for stereo depth recognition, reduce the data calculation amount of the entire depth recognition solution, improve the universality and recognition accuracy of the depth recognition solution, so as to ensure that the stereo depth recognition solution provided by this application can be effectively compatible with robots of different models / scenes to achieve high-precision and high-efficiency stereo depth recognition effects.

[0089] In this application, to ensure that the control device 10 can accurately measure the initial disparity map between the right-eye scene map and the left-eye scene map of the same robot, the embodiments of this application implement the foregoing function by providing an initial disparity map measurement solution. The following provides a detailed description of the initial disparity map measurement solution provided by this application.

[0090] Please refer to Figure 3 , Figure 3 is Figure 2 a schematic flowchart of the sub-steps included in step S220 in

[0091] Sub-step S221: Feature extraction is respectively performed on the right-eye scene map and the left-eye scene map to obtain corresponding right-eye scene feature maps and left-eye scene feature maps.

[0092] In this embodiment, the control device 10 can perform feature extraction on the right-eye scene map and the left-eye scene map respectively from the dimensions of bath (number of image frames), channels (number of image channels), height (number of pixels in the vertical direction of the image), and width (number of pixels in the horizontal direction of the image) to obtain right-eye scene feature maps that match the right-eye scene map and left-eye scene feature maps that match the left-eye scene map.

[0093] Sub-step S222: Pixel feature matching is performed between the left-eye scene feature map and the right-eye scene feature map to obtain the disparity value between each pixel point in the left-eye scene feature map and the pixel point with feature matching in the right-eye scene feature map.

[0094] In this embodiment, for each first pixel point in the left-eye scene feature map, the control device 10 can select each second pixel point in the right-eye scene feature map that is in the same horizontal direction of the image as this first pixel point and perform pixel feature matching with this first pixel point to determine the target second pixel point that successfully matches the feature of this first pixel point, and then determine the disparity value between this first pixel point and the target second pixel point. At this time, the disparity value between each pixel point in the left-eye scene feature map and the pixel point with feature matching in the right-eye scene feature map can be obtained. During this process, the control device 10 can implement the foregoing pixel feature matching function by calling the Concatenate function.

[0095] Sub-step S223: Image aggregation is performed on the disparity values corresponding to all pixel points in the left-eye scene feature map to obtain the initial disparity map.

[0096] In this embodiment, after determining the disparity values ​​corresponding to all pixel points in the left-eye scene feature map, the control device 10 can use the pixel position of each pixel point in the left-eye scene feature map as a reference template, perform image aggregation on the disparity values ​​of each pixel point, and obtain an initial disparity map.

[0097] Therefore, the present application can accurately measure the initial disparity map between the right eye scene map and the left eye scene map of the same robot by executing the above sub-steps S221 to S223.

[0098] In the present application, in order to ensure that the control device 10 can effectively perform the disparity range classification operation, the present application embodiment implements the above functions by providing a disparity range classification scheme. The disparity range classification scheme provided by the present application is described in detail below.

[0099] Please refer to Figure 4 , Figure 4 yes Figure 2 Schematic diagram of the flow of sub-steps included in step S230. In the embodiment of the present application, step S230 may include sub-steps S231 to S233.

[0100] Sub-step S231 , determining the disparity range specification corresponding to the initial disparity map according to the maximum disparity value in the initial disparity map and the preset disparity range planning step.

[0101] In this embodiment, the disparity range specification includes a plurality of disparity ranges that are continuously distributed, and the disparity range planning step length is the length of the interval maintained for each disparity range.

[0102] In sub-step S232 , a cross entropy loss function is called according to the disparity range specification to identify the disparity ranges to which all disparity values ​​in the initial disparity map belong.

[0103] Sub-step S233 , fitting a distribution probability curve of the disparity range at the initial disparity map according to the respective occurrence frequencies of different disparity ranges corresponding to the initial disparity map, to obtain a disparity range distribution result.

[0104] In this embodiment, the disparity range distribution result obtained by the curve fitting operation can be expressed in the form of a normal distribution curve. In this case, the disparity range distribution result describes the photographing capability of the current binocular camera of the target robot for the scene to be tested.

[0105] Therefore, the present application can perform an effective disparity range classification operation on the initial disparity map between the right eye scene map and the left eye scene map of the target robot by executing the above sub-steps S231 to S233.

[0106] In this application, to ensure that the control device 10 can migrate the parallax data of the target robot that directly matches the captured binocular scene map into a preset expected data domain, so as to reduce the data computation amount of the depth recognition solution and improve the universality and recognition accuracy of the corresponding depth recognition solution, the embodiments of this application implement the foregoing function through a parallax range distribution update solution. The following provides a detailed description of the parallax range distribution update solution provided by this application.

[0107] Please refer to Figure 5 , Figure 5 is Figure 2 a schematic flowchart of the sub-steps included in step S240 in

[0108] Sub-step S241: Determine whether the standard deviation value of the normal distribution curve corresponding to the parallax range distribution result is greater than a preset standard deviation threshold.

[0109] In this embodiment, the control device 10 can use the standard deviation value of the expected normal distribution as the preset standard deviation threshold, so that the control device 10 compares the standard deviation value of the normal distribution curve corresponding to the parallax range distribution result with the preset standard deviation threshold to determine whether the to-be-recognized image data corresponding to the parallax range distribution result is within the expected data domain. Among them, if the standard deviation value of the normal distribution curve corresponding to the parallax range distribution result is greater than the preset standard deviation threshold, it means that the to-be-recognized image data corresponding to the parallax range distribution result is not within the expected data domain. At this time, the control device 10 will correspondingly execute sub-step S242; if the standard deviation value of the normal distribution curve corresponding to the parallax range distribution result is less than or equal to the preset standard deviation threshold, it means that the to-be-recognized image data corresponding to the parallax range distribution result is already within the expected data domain. At this time, the control device 10 will correspondingly execute sub-step S243.

[0110] Sub-step S242: Calculate the mathematical expectation difference and standard deviation multiple between the expected normal distribution and the normal distribution curve, and perform a curve shift on the normal distribution curve corresponding to the parallax range distribution result according to the mathematical expectation difference and standard deviation multiple to obtain an optimized parallax distribution result.

[0111] In this embodiment, if the image data to be identified corresponding to the disparity range distribution result is not within the expected data domain, the control device 10 will calculate the difference between the mathematical expectation of the expected normal distribution and the mathematical expectation of the normal distribution curve (i.e., the mathematical expectation difference), and calculate the multiple value of the standard deviation value of the expected normal distribution relative to the standard deviation value of the normal distribution curve (i.e., the standard deviation multiple), and then the control device 10 will perform a curve shift on the normal distribution curve corresponding to the disparity range distribution result based on the mathematical expectation difference and the standard deviation multiple, so that the mathematical expectation of the normal distribution curve corresponding to the disparity distribution optimization result is consistent with the mathematical expectation of the expected normal distribution, and the standard deviation value of the normal distribution curve corresponding to the disparity distribution optimization result is consistent with the standard deviation value of the expected normal distribution.

[0112] Sub-step S243 , directly taking the disparity range distribution result as the disparity distribution optimization result.

[0113] In this embodiment, if the to-be-recognized image data corresponding to the disparity range distribution result is within the expected data domain, it means that the disparity distribution result does not need to be updated currently. In this case, the disparity range distribution result can be directly used as the disparity distribution optimization result.

[0114] Therefore, the present application can migrate the image data to be identified in different conditions to the same expected data domain for stereo depth recognition operations by executing the above sub-steps S241 to S243, thereby reducing the data calculation amount of the entire depth recognition scheme, and improving the universality and recognition accuracy of the depth recognition scheme, so as to ensure that the stereo depth recognition scheme provided by the present application can be effectively compatible with robots of different styles / scenes to achieve high-precision and high-efficiency stereo depth recognition effects.

[0115] In the present application, in order to ensure that the control device 10 can use the disparity value data migrated to the expected data domain to perform a high-precision disparity prediction operation and obtain a high-precision target scene disparity map of the target robot that matches the scene to be tested, the embodiment of the present application implements the above function by providing a scene disparity prediction scheme. The scene disparity prediction scheme provided by the present application is described in detail below.

[0116] Please refer to Figure 6 , Figure 6 yes Figure 2 Schematic diagram of the flow of sub-steps included in step S250. In the embodiment of the present application, step S250 may include sub-steps S251 to S252.

[0117] Sub-step S251: Based on the initial disparity map, call the disparity regression model to estimate all possible disparity values that satisfy the disparity distribution optimization result and the occurrence probability of each of all possible disparity values for each pixel in the target scene disparity map.

[0118] In this embodiment, the image size of the target scene disparity map is the same as that of the left-eye scene map. The control device 10 can use a disparity regression model trained based on the Softargmin algorithm and the smoothed L1 norm loss function to perform disparity prediction, and obtain all possible disparity values that satisfy the disparity distribution optimization result and the occurrence probability of each of all possible disparity values for each pixel in the target scene disparity map.

[0119] Sub-step S252: For each pixel in the target scene disparity map, perform a weighted summation operation on all possible disparity values corresponding to the pixel and the occurrence probability of each of all possible disparity values to obtain the target scene disparity value at the pixel in the target scene disparity map.

[0120] Thus, this application can perform high-precision disparity prediction operations by executing the above sub-step S251 to sub-step S252, using the disparity value data migrated into the expected data domain, and obtain a high-precision target scene disparity map of the target robot that matches the scene to be measured.

[0121] In this application, to ensure that the control device 10 can effectively perform the disparity depth conversion operation, this application embodiment realizes the foregoing function by providing a disparity depth conversion scheme. The disparity depth conversion scheme provided by this application is described in detail below.

[0122] Please refer to Figure 7 , Figure 7 is Figure 2 a schematic flowchart of the sub-steps included in step S260 in

[0123] Sub-step S261: For each pixel in the target scene disparity map, substitute the camera focal length, the camera baseline distance, and the target scene disparity value corresponding to the pixel into the pre-stored disparity-depth correlation relationship to calculate the target scene depth value corresponding to the pixel in the stereo depth distribution map.

[0124] In this embodiment, referring to Figure 8 the effect display diagram of the disparity-depth correlation relationship shown, P is used to represent the object to be photographed, and O L is used to represent the projection center of the left-eye camera of the target robot, and O Ra is used to represent the projection center of the right - eye camera of the target robot, b is used to represent the distance between the projection centers of the left - eye camera and the right - eye camera respectively (i.e., the camera baseline distance), f is used to represent the camera focal length provided by a single camera at the target robot, L is used to represent the number of image - formable pixels of the camera plane of the corresponding camera (including the left - eye camera and the right - eye camera) in the imaging horizontal direction, P L is used to represent the imaging point of the object to be photographed on the camera plane of the left - eye camera, P R is used to represent the imaging point of the object to be photographed on the camera plane of the right - eye camera, Z is used to represent the stereo depth value between the object to be photographed and the target robot, X L is used to represent the distance between the left - eye camera imaging point P L and the left imaging boundary of the left - eye camera, X R is used to represent the distance between the right - eye camera imaging point P R and the left imaging boundary of the left - and right - eye cameras. The parallax value of the object to be photographed between the left - eye camera and the right - eye camera is |X L - X R |. At this time, the parallax - depth correlation relationship can be expressed by the following formula:

[0125]

[0126] where D is used to represent the parallax value of the object to be photographed between the left - eye camera and the right - eye camera of the target robot on the same baseline, b is used to represent the camera baseline distance between the left - eye camera and the right - eye camera of the target robot, f is used to represent the camera focal length of the target robot, and Z is used to represent the stereo depth value between the object to be photographed and the target robot.

[0127] Thus, the control device 10 can calculate the target - scene depth value matched by the target - scene parallax value of each pixel point in the target - scene parallax map based on the above - mentioned parallax - depth correlation relationship.

[0128] Sub - step S262: Aggregate the target - scene depth values corresponding to all pixel points in the target - scene parallax map to obtain a stereo depth distribution map.

[0129] In this embodiment, after the control device 10 determines the target - scene depth value corresponding to each pixel point in the target - scene parallax map, it can use the pixel positions of the pixel points in the target - scene parallax map as a reference template to aggregate all the target - scene depth values to obtain the stereo depth distribution map.

[0130] Thus, by executing the above sub-steps S261 and S262, the present application can perform an effectively substantial parallax depth conversion operation on the target scene parallax map with high precision, achieving a high-precision and high-efficiency stereo depth recognition effect.

[0131] In the present application, to ensure that the control device 10 can execute the above stereo depth recognition method through the stereo depth recognition device 100, the present application realizes the foregoing functions by means of functional module division of the stereo depth recognition device 100. The following describes the specific composition of the stereo depth recognition device 100 provided by the present application.

[0132] Please refer to Figure 9 , Figure 9 which is a schematic diagram of the composition of the stereo depth recognition device 100 provided by an embodiment of the present application. In the embodiment of the present application, the stereo depth recognition device 100 may include a scene view acquisition module 110, an initial parallax determination module 120, a parallax distribution classification module 130, a parallax distribution optimization module 140, a scene parallax regression module 150, and a stereo depth calculation module 160.

[0133] The scene view acquisition module 110 is configured to acquire a left-eye scene map and a right-eye scene map captured by the target robot for the scene to be measured.

[0134] The initial parallax determination module 120 is configured to determine an initial parallax map of the right-eye scene map relative to the left-eye scene map.

[0135] The parallax distribution classification module 130 is configured to classify the parallax values in the initial parallax map into parallax range classifications to obtain a corresponding parallax range distribution result.

[0136] The parallax distribution optimization module 140 is configured to update the parallax range distribution result of the initial parallax map according to a preset desired normal distribution to obtain a corresponding parallax distribution optimization result.

[0137] The scene parallax regression module 150 is configured to call a pre-stored parallax regression model according to the parallax distribution optimization result to perform parallax regression processing on the initial parallax map to obtain a target scene parallax map matching the scene to be measured.

[0138] The stereo depth calculation module 160 is configured to perform parallax depth conversion processing on the target scene parallax map according to the camera focal length and the camera baseline distance of the target robot to obtain a stereo depth distribution map between the target robot and the scene to be measured.

[0139] It should be noted that for the three-dimensional depth recognition device 100 provided in the embodiments of the present application, its basic principle and the resulting technical effects are the same as those of the aforementioned three-dimensional depth recognition method. For a brief description, for the parts not mentioned in this embodiment, reference may be made to the description of the three-dimensional depth recognition method above.

[0140] In the embodiments provided in the present application, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the device, method, and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0141] In addition, in each embodiment of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. If the functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned readable storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0142] In summary, in a three-dimensional depth recognition method, device, control device, and readable storage medium provided by the present application, after determining an initial disparity map between a left-eye scene map and a right-eye scene map captured by a target robot for a to-be-detected scene, the present application can classify the disparity values in the initial disparity map into disparity range categories to obtain a disparity range distribution result, and update the disparity range distribution result according to a preset expected normal distribution to obtain an optimized disparity distribution result. Then, according to the optimized disparity distribution result, a pre-stored disparity regression model is called to perform disparity regression processing on the initial disparity map to obtain a target scene disparity map matching the to-be-detected scene. Next, according to the camera focal length and camera baseline distance of the target robot, disparity depth conversion processing is performed on the target scene disparity map to obtain a three-dimensional depth distribution map between the target robot and the to-be-detected scene. Thus, through the disparity range distribution update operation, image data to be recognized in different situations is migrated to the same expected data domain for three-dimensional depth recognition, reducing the data calculation amount of the entire depth recognition solution, improving the universality and recognition accuracy of the depth recognition solution, and ensuring that the three-dimensional depth recognition solution provided by the present application can be effectively compatible with robots of different models / scenes to achieve a high-precision and high-efficiency three-dimensional depth recognition effect.

[0143] As described above, the above are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A three-dimensional depth recognition method, characterized in that, The method includes: Obtaining a left-eye scene image and a right-eye scene image captured by a target robot for a scene to be measured; Determining an initial disparity map of the right-eye scene image relative to the left-eye scene image; Classifying each disparity value in the initial disparity map into disparity range categories to obtain a corresponding disparity range distribution result; Updating the disparity range distribution result of the initial disparity map according to a preset expected normal distribution to obtain a corresponding optimized disparity distribution result; wherein, the step of updating the disparity range distribution result of the initial disparity map according to a preset expected normal distribution to obtain a corresponding optimized disparity distribution result includes: judging whether the standard deviation value of the normal distribution curve corresponding to the disparity range distribution result is greater than a preset standard deviation threshold; if it is determined that the standard deviation value is greater than the preset standard deviation threshold, calculating the mathematical expectation difference and standard deviation multiple between the expected normal distribution and the normal distribution curve, and performing curve offset on the normal distribution curve corresponding to the disparity range distribution result according to the mathematical expectation difference and the standard deviation multiple to obtain the optimized disparity distribution result; if it is determined that the standard deviation value is less than or equal to the preset standard deviation threshold, directly using the disparity range distribution result as the optimized disparity distribution result; Invoking a pre-stored disparity regression model according to the optimized disparity distribution result to perform disparity regression processing on the initial disparity map to obtain a target scene disparity map matching the scene to be measured; wherein, the step of invoking a pre-stored disparity regression model according to the optimized disparity distribution result to perform disparity regression processing on the initial disparity map to obtain a target scene disparity map matching the scene to be measured includes: on the basis of the initial disparity map, invoking the disparity regression model to estimate all possible disparity values that satisfy the optimized disparity distribution result and the occurrence probabilities of all possible disparity values for each pixel point in the target scene disparity map; for each pixel point in the target scene disparity map, performing a weighted summation operation on all possible disparity values corresponding to the pixel point and the occurrence probabilities of all possible disparity values to obtain the target scene disparity value at the pixel point in the target scene disparity map; Performing disparity-depth conversion processing on the target scene disparity map according to the camera focal length and camera baseline distance of the target robot to obtain a stereo depth distribution map between the target robot and the scene to be measured.

2. The method according to claim 1, characterized in that, The step of determining an initial disparity map of the right-eye scene image relative to the left-eye scene image includes: Performing feature extraction on the right-eye scene image and the left-eye scene image respectively to obtain corresponding right-eye scene feature maps and left-eye scene feature maps; Performing pixel feature matching between the left-eye scene feature map and the right-eye scene feature map to obtain the disparity values between each pixel point in the left-eye scene feature map and the pixel points with feature matching in the right-eye scene feature map; Performing image aggregation on the disparity values corresponding to all pixel points in the left-eye scene feature map to obtain the initial disparity map.

3. The method according to claim 1, characterized in that, The step of classifying the disparity values in the initial disparity map into disparity range categories to obtain the corresponding disparity range distribution result includes: Determining the disparity range specification corresponding to the initial disparity map according to the maximum disparity value in the initial disparity map and a preset step size for disparity range planning; Invoking a cross-entropy loss function according to the disparity range specification to identify the disparity ranges to which all the disparity values in the initial disparity map belong; Fitting a distribution probability curve of the disparity ranges at the initial disparity map according to the occurrence frequencies of the different disparity ranges corresponding to the initial disparity map to obtain the disparity range distribution result.

4. The method according to any one of claims 1 to 3, characterized in that The step of performing disparity-depth conversion processing on the target scene disparity map according to the camera focal length and camera baseline distance of the target robot to obtain a stereo depth distribution map between the target robot and the to-be-measured scene includes: For each pixel point in the target scene disparity map, substituting the camera focal length, the camera baseline distance, and the disparity value of the target scene corresponding to this pixel point into a pre-stored disparity-depth correlation relationship to calculate the target scene depth value corresponding to this pixel point in the stereo depth distribution map; Performing image aggregation on the target scene depth values corresponding to all the pixel points in the target scene disparity map to obtain the stereo depth distribution map.

5. The method according to claim 4, characterized in that, The disparity-depth correlation relationship is expressed by the following formula: where D is used to represent the disparity value between the left-eye camera and the right-eye camera of the target robot on the same baseline for the object being photographed, b is used to represent the camera baseline distance between the left-eye camera and the right-eye camera of the target robot, f is used to represent the camera focal length of the target robot, and Z is used to represent the stereo depth value between the object being photographed and the target robot.

6. A three-dimensional depth recognition device, characterized in that, The device includes: A scene view acquisition module, configured to acquire a left-eye scene map and a right-eye scene map captured by the target robot for the to-be-measured scene; An initial disparity determination module, configured to determine an initial disparity map of the right-eye scene map relative to the left-eye scene map; A disparity distribution classification module, configured to classify the disparity values in the initial disparity map into disparity range categories to obtain the corresponding disparity range distribution result; A disparity distribution optimization module, configured to update the disparity range distribution result of the initial disparity map according to a preset expected normal distribution to obtain the corresponding disparity distribution optimization result; wherein, the disparity distribution optimization module is specifically configured to: determine whether the standard deviation value of the normal distribution curve corresponding to the disparity range distribution result is greater than a preset standard deviation threshold; if it is determined that the standard deviation value is greater than the preset standard deviation threshold, calculate the mathematical expectation difference and standard deviation multiple between the expected normal distribution and the normal distribution curve, and perform curve offset on the normal distribution curve corresponding to the disparity range distribution result according to the mathematical expectation difference and the standard deviation multiple to obtain the disparity distribution optimization result; if it is determined that the standard deviation value is less than or equal to the preset standard deviation threshold, directly use the disparity range distribution result as the disparity distribution optimization result; A scene parallax regression module, which is configured to call a pre-stored parallax regression model according to the optimized result of the parallax distribution, perform parallax regression processing on the initial parallax map, and obtain a target scene parallax map matching the to-be-detected scene; wherein, the scene parallax regression module is specifically configured to: on the basis of the initial parallax map, call the parallax regression model to estimate all possible parallax values and the occurrence probabilities of all possible parallax values that satisfy the optimized result of the parallax distribution for each pixel point in the target scene parallax map; for each pixel point in the target scene parallax map, perform a weighted summation operation on all possible parallax values corresponding to the pixel point and the occurrence probabilities of all possible parallax values, and obtain the target scene parallax value at the pixel point of the target scene parallax map. A stereo depth calculation module, which is configured to perform parallax depth conversion processing on the target scene parallax map according to the camera focal length and the camera baseline distance of the target robot, and obtain a stereo depth distribution map between the target robot and the to-be-detected scene.

7. A control device, characterized in that, It includes a processor and a memory, the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the stereo depth recognition method according to any one of claims 1-5.

8. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the stereo depth recognition method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Machine vision depth estimation method, device and system

    CN113034568A

  • Apparatus and method for performing 3D estimation based on locally determined 3D information hypotheses

    EP3252713A1