Binocular ranging method, device and storage medium based on neural network
By annotating and matching the targets of the binocular image, and using neural network encoder and decoder for depth estimation, the problem of large amount of calculation in the prior art is solved, and efficient and accurate binocular image matching is achieved.
Patent Information
- Application Number
- CN202111201934.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-15
AI Technical Summary
In the prior art, the calculation of binocular image matching based on neural networks is large and affects efficiency.
By annotating and synchronously calibrating multiple targets of the left-eye image and the right-eye image, the target sequence information is output using the neural network encoder, and target matching and depth estimation are performed in the decoder, and the target depth estimation error is calculated to train the neural network.
Effectively reduce the amount of calculation, improve the efficiency of binocular image matching, ensure matching accuracy, and accurately estimate depth values, especially in long distances and textureless areas.
Abstract
Description
Technical Field
[0001] The present invention relates to a binocular ranging method, device and storage medium based on a neural network. Background Art
[0002] In the prior art, when using a neural network to perform matching training on binocular images, based on full-image matching, the amount of calculation is large, which affects the matching efficiency. Summary of the Invention
[0003] The object of the present invention is to provide a binocular ranging method, device and storage medium based on a neural network, which can effectively improve the matching efficiency of binocular images.
[0004] Based on the same inventive concept, the present invention has three independent technical solutions:
[0005] 1. A binocular ranging method based on a neural network, characterized by comprising the following steps:
[0006] Step 1: Label multiple first targets in the left-eye image and the right-eye image. Each first target contains first target sequence information. Synchronize and calibrate the left-eye image and the right-eye image, and input them into a neural network encoder;
[0007] Step 2: For the input left-eye image and right-eye image, the neural network encoder correspondingly outputs multiple second targets, and each second target contains second target sequence information;
[0008] Step 3: Input the multiple first target sequences and multiple second target sequences into a neural network decoder, perform target matching on the multiple first targets and the multiple second targets, and calculate the target depth estimation error.
[0009] Further, the first target sequence includes: target category, target rectangle box, feature vector of the target, target indirect depth value, length, width and height of the target, and labeled feature points of the target.
[0010] Further, the second target sequence includes: target existence probability, target category probability, target rectangle box, target indirect depth value, length, width and height of the target, and associated feature points of the target.
[0011] Further, in step 3, performing target matching on the multiple first targets and the multiple second targets is achieved by the following method
[0012] Step 3.1: Arrange the multiple second targets in descending order according to the target existence probability;
[0013] Step 3.2: Match the multiple first targets with the sorted multiple second targets one by one. For each first target, use the second target with the largest intersection over union (IoU) of the target rectangular bounding boxes as the matching target for this first target.
[0014] Further, it includes Step 3.3: Calculate the matching cost for each matching target. The matching cost includes the distance error of the center points of the rectangular bounding boxes, the length-width-height error of the rectangular bounding boxes, the distance error of the feature points, the length-width-height error of the targets, and the indirect depth error.
[0015] Further, in Step 3, perform target matching on the multiple first targets and the multiple second targets, which is achieved by the following method.
[0016] For the first target and the second target, calculate the cosine similarity of the target feature vectors; for each first target, use the second target with the largest cosine similarity as the matching target for this first target.
[0017] Further, in Step 3, calculate the target depth estimation error by the following method.
[0018] For the first target, calculate the true depth of the feature points according to the feature point positions and the target indirect depth values.
[0019] For the matching target of this first target, match the feature points of the first target with the feature points of the matching target in order, and calculate the estimated depth of the feature points of the matching target.
[0020] Calculate the difference between the true depth and the estimated depth.
[0021] Further, weight the difference between the true depth and the estimated depth to obtain the depth loss of the neural network; train the neural network according to the depth loss.
[0022] 2. A binocular ranging device based on a neural network, which is used to execute the above method.
[0023] 3. A storage medium, on which a computer program is stored. When the program is executed by a processor, the above method is implemented.
[0024] The beneficial effects of the present invention are as follows:
[0025] The present invention only estimates the depth of the targets in the binocular images, and uses the target feature points to ensure the estimated depth through constraints, without full-image matching, which greatly reduces the amount of calculation, effectively improves the matching efficiency of the binocular images based on the neural network, and ensures the matching accuracy.
[0026] When the present invention estimates the depth of a target, for a long distance where it is impossible to effectively collect ground truth, a reliable depth value can also be obtained through binocular constraints; for the case of only a monocular image, the present invention can also stably estimate a depth value. The present invention can effectively avoid the situation of misestimating the depth when there is no texture or weak texture in the target area, and ensure the matching accuracy. Specific embodiments
[0027] The present invention will be described in detail below in conjunction with each embodiment. However, it should be noted that these embodiments are not intended to limit the present invention, and any equivalent transformation or substitution in function, method, or structure made by those of ordinary skill in the art according to these embodiments shall fall within the protection scope of the present invention.
[0028] Embodiment 1:
[0029] Binocular ranging method based on neural network
[0030] The method includes the following steps:
[0031] Step 1: Label multiple first targets in the left-eye image and the right-eye image. Each first target contains first target sequence information. Synchronize and calibrate the left-eye image and the right-eye image, and input them into the neural network encoder.
[0032] The first target sequence includes: target category, target rectangle box, feature vector of the target, indirect depth value of the target, length, width and height of the target, and labeled feature points of the target.
[0033] Step 2: For the input left-eye image and right-eye image, the neural network encoder correspondingly outputs multiple second targets, and each second target contains second target sequence information.
[0034] The second target sequence includes: target existence probability, target category probability, target rectangle box, indirect depth value of the target, length, width and height of the target, and associated feature points of the target.
[0035] Step 3: Input the multiple first target sequences and multiple second target sequences into the neural network decoder, perform target matching on the multiple first targets and multiple second targets, and calculate the target depth estimation error.
[0036] In step 3, the target matching of the multiple first targets and multiple second targets is realized by the following method:
[0037] Step 3.1: Arrange the multiple second targets in descending order according to the target existence probability;
[0038] Step 3.2: Match the multiple first targets with the sorted multiple second targets one by one. For each first target, use the second target with the largest intersection over union (IoU) of the target rectangular boxes as the matching target for this first target.
[0039] Step 3.3: Calculate the matching cost of each matching target. The matching cost includes the distance error of the center points of the rectangular boxes, the length, width, and height errors of the rectangular boxes, the distance error of the feature points, the length, width, and height errors of the targets, and the indirect depth error.
[0040] In Step 3, to perform target matching between the multiple first targets and the multiple second targets, it can also be achieved by the following method:
[0041] For the first target and the second target, calculate the cosine similarity of the target feature vectors; for each first target, use the second target with the largest cosine similarity as the matching target for this first target.
[0042] In Step 3, calculate the target depth estimation error by the following method:
[0043] For the first target, calculate the true depth of the feature points according to the feature point positions and the target indirect depth values;
[0044] For the matching target of this first target, match the feature points of the first target with the feature points of the matching target in order, and calculate the estimated depth of the feature points of the matching target;
[0045] Calculate the difference between the true depth and the estimated depth.
[0046] Based on the difference between the true depth and the estimated depth, obtain the depth loss of the neural network by weighting; according to the depth loss, train the neural network.
[0047] Embodiment 2:
[0048] Binocular ranging device based on neural network
[0049] The binocular ranging device is used to execute the method described in Embodiment 1.
[0050] Embodiment 3:
[0051] Storage medium
[0052] A computer program is stored on the storage medium, and when the program is executed by a processor, it implements the method described in Embodiment 1.
[0053] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present invention, and they are not intended to limit the protection scope of the present invention. Any equivalent implementation manners or modifications made without departing from the technical spirit of the present invention should be included within the protection scope of the present invention.
[0054] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention.
Claims
1. A method for calculating the depth of a binocular image target based on a neural network, characterized in that, It includes the following steps: Step 1: Label multiple first targets in the left-eye image and the right-eye image. Each first target contains first target sequence information. Synchronize and calibrate the left-eye image and the right-eye image, and input them into the neural network encoder; Step 2: For the input left-eye image and right-eye image, the neural network encoder correspondingly outputs multiple second targets. Each second target contains second target sequence information; Step 3: Perform target matching on the multiple first targets and the multiple second targets, and calculate the target depth estimation error; The first target sequence information includes: target category, target rectangle box, feature vector of the target, target indirect depth value, length, width and height of the target, labeled feature points of the target; The second target sequence information includes: target existence probability, target category probability, target rectangle box, target indirect depth value, length, width and height of the target, associated feature points of the target; In Step 3, the target matching of the multiple first targets and the multiple second targets is realized by the following method Step 3.1: Arrange the multiple second targets in descending order according to the target existence probability; Step 3.2: Match the multiple first targets with the sorted multiple second targets one by one. For each first target, take the second target with the largest intersection over union of the target rectangle boxes as the matching target of this first target; In Step 3, the target depth estimation error is calculated by the following method For the first target, calculate the true depth of the labeled feature points according to the position of the labeled feature points of the target and the target indirect depth value; For the matching target of this first target, match the labeled feature points of the first target with the associated feature points of the matching target in order, and calculate the estimated depth of the associated feature points of the matching target; Calculate the difference between the true depth and the estimated depth.
2. The binocular image target depth calculation method based on a neural network according to claim 1, characterized in that: According to the difference between the true depth and the estimated depth, weight to obtain the depth loss of the neural network; train the neural network according to the depth loss.
3. The binocular image target depth calculation method based on a neural network according to claim 1, characterized in that: In Step 3, it further includes Step 3.3: Calculate the matching cost of each matching target. The matching cost includes the center point distance error of the rectangle box, the length, width and height error of the rectangle box, the distance error of the feature points, the length, width and height error of the target, and the indirect depth error.
4. A storage medium has a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Depth information determination method and related device
CN108537837A
Deep estimation network training method and device for automatic driving scene and autonomous vehicle
CN111428859A