A binocular image matching method, device, equipment and storage medium

By detecting and determining bounding boxes in binocular images and using the bounding boxes to regress the target location, the problem of complex matching algorithms in the prior art is solved, and efficient and accurate target matching is achieved.

CN115690469BActive Publication Date: 2026-02-03BEIJING TUSEN ZHITU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110873191.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-30
Publication Date
2026-02-03
Estimated Expiration
2041-07-30

AI Technical Summary

Technical Problem

In existing technologies, binocular image matching algorithms are complex and have a high matching error rate, which affects the robustness and accuracy of target matching.

Method used

By performing target detection on the first image of the binocular image, a first bounding box is obtained, and the corresponding second bounding box is determined in the second image. The third bounding box of the target in the second image is then regressed using the second bounding box, which reduces computational overhead and improves matching efficiency and accuracy.

Benefits of technology

It achieves accurate target matching within binocular images, reduces computational overhead, and improves the efficiency, accuracy, and robustness of matching, while avoiding the limitations of complex matching algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690469B_ABST
    Figure CN115690469B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a binocular image matching method, device, equipment and storage medium. The method comprises: performing target detection on a first image of a binocular image to obtain a first bounding box of a target in the first image; determining a second bounding box corresponding to the first bounding box in a second image of the binocular image; and regressing a third bounding box of the target in the second image by using the second bounding box. The technical scheme provided by the embodiments of the present application realizes accurate matching of the target in the binocular image, without performing target detection on both images in the binocular image, and then matching the detected targets in the two images by using a matching algorithm, thereby greatly reducing the calculation cost of the target matching in the binocular image, avoiding the limitation of the target matching in the binocular image, and improving the efficiency, accuracy and robustness of the target matching in the binocular image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image data processing technology, and in particular to a binocular image matching method, apparatus, device and storage medium. Background Technology

[0002] Target ranging is a crucial part of autonomous driving systems. To ensure the accuracy of target ranging, binocular ranging is usually used. The corresponding targets are detected separately from the binocular images, and then the targets in the binocular images are matched to determine the position of the same target in the binocular images. Then, the distance of the target is calculated based on parallax or triangulation by combining the intrinsic and extrinsic parameters of the binocular camera.

[0003] Currently, when performing target matching on binocular images, corresponding matching algorithms are usually pre-designed to establish the correspondence between the same target in the binocular images. At this time, the matching algorithm involves a variety of additional features, multiple thresholds, and fine parameter adjustments, which makes the matching algorithm more complex and has a certain matching error rate, greatly affecting the robustness and accuracy of target matching in binocular images. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a binocular image matching method, apparatus, device, and storage medium to achieve accurate matching of targets within binocular images and improve the accuracy and robustness of target matching within binocular images.

[0005] In a first aspect, embodiments of the present invention provide a binocular image matching method, the method comprising:

[0006] Target detection is performed on the first image of the binocular image to obtain the first bounding box of the target in the first image;

[0007] In the second image of the binocular image, determine the second bounding box corresponding to the first bounding box;

[0008] The third bounding box of the target within the second image is regressed using the second bounding box.

[0009] In a second aspect, embodiments of the present invention provide a binocular image matching device, the device comprising:

[0010] The target detection module is used to perform target detection on the first image and obtain the first bounding box of the target in the first image;

[0011] The region determination module is used to determine the second bounding box corresponding to the first bounding box in the second image of the stereo image;

[0012] A regression module is used to regress a third bounding box of the target within the second image using the second bounding box.

[0013] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising:

[0014] One or more processors;

[0015] Storage device for storing one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the binocular image matching method according to any embodiment of the present invention.

[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the binocular image matching method described in any embodiment of the present invention.

[0018] This invention provides a binocular image matching method, apparatus, device, and storage medium. By performing target detection on the first image of a binocular image, a first bounding box of the target within the first image can be obtained. At this point, the second bounding box corresponding to the first bounding box can be directly determined in the second image of the binocular image. Then, the third bounding box of the target within the second image is regressed using the second bounding box, thereby achieving accurate target matching within the binocular image. This eliminates the need to perform target detection on both images of the binocular image and then use a matching algorithm to match the targets detected in the two images, greatly reducing the computational overhead of target matching within the binocular image. It solves the problem of the complexity of matching algorithms using multiple features and thresholds in the prior art, avoids the limitations of target matching within the binocular image, and improves the efficiency, accuracy, and robustness of target matching within the binocular image. Attached Figure Description

[0019] Figure 1A This is a flowchart of a binocular image matching method provided in Embodiment 1 of the present invention;

[0020] Figure 1B This is a schematic diagram illustrating the determination of the second bounding box in the method provided in Embodiment 1 of the present invention;

[0021] Figure 1C This is a schematic diagram illustrating the principle of target imaging in a binocular camera in the method provided in Embodiment 1 of the present invention;

[0022] Figure 2A This is a flowchart of a binocular image matching method provided in Embodiment 2 of the present invention;

[0023] Figure 2BThis is a schematic diagram illustrating the principle of the binocular image matching process provided in Embodiment 2 of the present invention;

[0024] Figure 3 This is a flowchart of a binocular image matching method provided in Embodiment 3 of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of a binocular image matching device provided in Embodiment 4 of the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present invention. Detailed Implementation

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the drawings, not all structures. Moreover, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0028] Example 1

[0029] Figure 1A This is a flowchart of a binocular image matching method provided in Embodiment 1 of the present invention. This embodiment is applicable to matching the same target existing in two binocular images. The binocular image matching method provided in this embodiment can be executed by the binocular image matching device provided in this embodiment of the present invention. This device can be implemented by software and / or hardware and integrated into the electronic device executing this method.

[0030] For details, please refer to Figure 1A The method may include the following steps:

[0031] S110, Perform target detection on the first image of the binocular image to obtain the first bounding box of the target in the first image.

[0032] Optionally, considering that when performing binocular ranging, it is necessary to match the same targets in the binocular images, but when performing target detection on the two images of the binocular images separately and then using the set matching algorithm to match each target in the two images, there will be a large matching calculation overhead, resulting in inefficient matching.

[0033] Therefore, to solve the above problems, this embodiment can perform target detection on the first image in the binocular image to obtain the position information of each target within the first image. Furthermore, to accurately mark each target within the first image, this embodiment can use bounding boxes (Bboxes) to represent the position detection results of each target within the first image, thus obtaining the first bounding boxes of each target within the first image. Subsequently, using the first bounding boxes of each target in the first image as a reference, and by analyzing the characteristics of the binocular camera used to acquire the binocular image, the position information of the same target in the second image can be determined, thereby achieving target matching within the binocular image.

[0034] Of course, this disclosure can also be applied to other multi-view images such as tri-view and quad-view images. In this case, any two views form a binocular image, and the position of the bounding box in other images can be determined based on the bounding box of any other image.

[0035] In this embodiment, the first image can be either the left or right image within the binocular image, and the other image is the second image of the binocular image. This embodiment does not limit whether the first or second image of the binocular image is the left or right image. Furthermore, the bounding box is a rectangular box used to outline each target within the first image. In this embodiment, the first bounding box can be represented as... Where N is the number of targets detected in the first image. This represents the first bounding box of the i-th target within the first image.

[0036] S120, determine the second bounding box corresponding to the first bounding box in the second image of the stereo image.

[0037] Specifically, after detecting the first bounding boxes of each target within the first image of the stereo camera, to reduce the computational overhead of target matching within the stereo image, target detection is not performed on the second image. Instead, using the intrinsic and extrinsic parameters of the stereo camera and referencing the first bounding boxes of each target in the first image, the second bounding boxes corresponding to the first bounding boxes of each target are directly determined in the second image. These second bounding boxes can be represented as follows: Used to mark the location of each target after preliminary matching in the second image.

[0038] For example, such as Figure 1B As shown, the second bounding box in this embodiment can be any of the following candidate boxes:

[0039] 1) A box at the same image position as the first bounding box is denoted as the first candidate box.

[0040] In the first scenario, when the imaging difference between the first and second images is small, it indicates that the positional difference of the same target within the first and second images is not significant. Therefore, in this embodiment, based on the coordinates of the first bounding boxes of each target in the first image, a box at the same position as the first bounding box can be directly determined in the second image and denoted as the first candidate box. Then, considering the small imaging difference between the first and second images, this first candidate box can be used as the second bounding box of each target in the second image.

[0041] It should be noted that this embodiment can use the parallax in binocular images to determine the imaging difference between the first image and the second image. This parallax is the difference in the imaging position of the target in the first image and the second image when the binocular camera captures the same target. For example... Figure 1C As shown, if point P is the target, O1 and O2 are the two optical centers of the stereo camera, the two line segments represent the virtual imaging planes corresponding to the two optical centers, f represents the focal length of the stereo camera, the baseline length of the two optical centers is called baseline, and the depth of the target point P is called depth. Therefore, P, O1, and O2 can form a triangle. Then, the disparity of point P on the virtual imaging planes corresponding to the two optical centers can be expressed as disparity = x. l -x r , where x l Let x be the position of point P in the first virtual imaging plane. r Let P be the position of point P in the second virtual imaging plane. Therefore, using the principle of similar triangles, we can obtain:

[0042] Considering that the depth of the same target is the same in both the first and second images, it can be determined from the above formula that the parallax in a binocular image is affected by the baseline and focal length of the binocular camera used to acquire the image.

[0043] Therefore, this embodiment can represent the situation where the imaging difference between the first image and the second image is small by pre-setting corresponding low disparity conditions. Furthermore, this embodiment can obtain the information of the binocular camera baseline and focal length corresponding to the currently matched binocular image, and then determine whether the binocular camera baseline and focal length meet the preset low disparity conditions, thereby determining whether the imaging difference between the first image and the second image is small. When the binocular camera baseline and focal length meet the preset low disparity conditions, the second bounding box in this embodiment is the first candidate box at the same image position as the first bounding box within the second image, without needing to adjust the first bounding box, ensuring the high efficiency of target matching within the binocular image. The low disparity conditions include, for example, the binocular camera baseline being less than a first threshold and / or the focal length being less than a second threshold. Of course, this is not limited to these; those skilled in the art can set low disparity conditions as needed, and this disclosure does not impose any limitations on this.

[0044] 2) The box obtained by translating the first bounding box and estimating the disparity is denoted as the second candidate box.

[0045] In the second scenario, due to the inevitable imaging differences between the first and second images, the position of the same target will also differ. Furthermore, from... Figure 1C As can be seen, the positional difference of the same target within a binocular image can be represented by disparity. Therefore, this embodiment can estimate the disparity of the binocular image, and then translate the first bounding box in the second image according to the estimated disparity to obtain the second bounding box of the same target in the second image. In other words, the second bounding box in this embodiment can be the second candidate box obtained by translating the first bounding box in the second image after estimating the disparity, thus ensuring the accuracy of target matching within the binocular image.

[0046] 3) The bounding box obtained by binocularly correcting the first bounding box and projecting it back to the uncorrected second image coordinate system is denoted as the third candidate bounding box. The second image coordinate system can be the image coordinate system of the second image, preferably the pixel coordinate system of the second image, but is not limited to this. The image coordinate system has its origin at the image center point, and the pixel coordinate system has its origin at the top-left corner of the image.

[0047] In the third scenario, the matching of the first and second bounding boxes of the same target within the binocular image is obtained with reference to an ideal binocular system. However, since the positions of the two optical centers in a binocular camera inevitably have some errors, it is difficult to obtain an ideal binocular system. Therefore, binocular correction is required to achieve an approximate ideal binocular system. In this embodiment, considering that the matching of targets within the binocular image is obtained by transforming the target bounding box and does not involve the transformation of other pixels within the binocular image, when determining the second bounding box of a target in the second image, the first bounding box of the target can be binocularly corrected and then projected back into the uncorrected second image coordinate system. This achieves the correction of target matching within the binocular image without needing to perform binocular correction on every pixel within the binocular image, greatly reducing the correction overhead of the binocular image. Furthermore, in this embodiment, the second bounding box can be the third candidate box obtained after binocularly correcting the first bounding box and projecting it back into the uncorrected second image coordinate system, ensuring the accuracy and efficiency of target matching within the binocular image.

[0048] 4) The first bounding box is subjected to binocular correction and translation to estimate disparity, and then projected back to the uncorrected second image coordinate system. The resulting box is denoted as the fourth candidate box.

[0049] In the fourth scenario, considering the inevitable imaging differences between the first and second images, and the limitations of an ideal binocular system, the second and third scenarios can be analyzed together. That is, the first bounding box undergoes binocular correction and translational parallax estimation, and is then projected back to the uncorrected second image coordinate system. This resulting box is denoted as the fourth candidate box, allowing the second bounding box in this embodiment to serve as the fourth candidate box, further improving the accuracy of target matching within the binocular image. Those skilled in the art should understand that translational parallax estimation can occur before or after binocular correction, and before or after projection. This disclosure does not limit the order of the translation operations. Furthermore, the uncorrected second image coordinate system mentioned in this disclosure can be either the image coordinate system of the second image or the pixel coordinate system of the second image; this disclosure does not impose any limitations on this.

[0050] Based on this, step S120 may include: translating the first bounding box by estimating the disparity to obtain a second bounding box at the corresponding position in the second image. Further, translating the first bounding box by estimating the disparity to obtain a second bounding box at the corresponding position in the second image includes: performing binocular correction on the first bounding box, translating by estimating the disparity, and projecting it back into the uncorrected second image coordinate system to obtain a second bounding box at the corresponding position in the second image.

[0051] It should be noted that the binocular correction and projection operations performed in this embodiment are all operations performed on the bounding box, and no correction operation is performed on the binocular image. This greatly saves the process of adjusting each pixel in the binocular image during binocular image correction, thereby improving the target matching efficiency of the binocular image.

[0052] S130, use the second bounding box to regress the third bounding box of the target in the second image.

[0053] In this embodiment, since the same target in the first and second images may have a certain size difference after imaging, and the second bounding box and the first bounding box are almost the same size or have little difference, it means that the second bounding box may not be able to fully enclose the corresponding target in the second image. Therefore, in order to ensure the accuracy of the target in the second image, after determining the second bounding box of the target in the second image, regression processing is performed on the second bounding box based on the target features in the second image to obtain the third bounding box of the target in the second image. The bounding box regression in this embodiment mainly analyzes a certain mapping relationship between the second bounding box and the actual features of the target in the second image, and uses this mapping relationship to map the second bounding box so that the obtained third bounding box is infinitely close to the true bounding box of the target, thereby achieving accurate target matching in the binocular images.

[0054] The target matching method within binocular images in this embodiment does not limit the application scenario and acquisition method of binocular images, avoiding the limitations of target matching within binocular images and improving the efficiency and robustness of target matching within binocular images. Furthermore, the binocular ranging disclosed herein is at the object level, not the pixel level, and can directly provide the object's depth information.

[0055] Furthermore, after using the second bounding box to regress the third bounding box of the target in the second image to achieve accurate matching of the target in the binocular image, this embodiment may further include: calculating the actual disparity of the target in the binocular image based on the first bounding box and the third bounding box; and calculating the depth information of the target based on the actual disparity, the binocular camera baseline and focal length corresponding to the binocular image.

[0056] In other words, after obtaining the first bounding box of the same target in the first image and the third bounding box in the second image within the stereo image, the coordinate difference between the first and second bounding boxes can be calculated as the actual disparity of the target in the stereo image. Then, the actual disparity of the target in the stereo image, the stereo camera baseline, and the focal length corresponding to the stereo image can be substituted into the formula. In this way, the depth information of the target can be calculated, thereby realizing target ranging within the binocular image.

[0057] The technical solution provided in this embodiment, by performing target detection on the first image of the binocular image, can obtain the first bounding box of the target in the first image. At this time, the second bounding box corresponding to the first bounding box can be directly determined in the second image of the binocular image. Then, the third bounding box of the target in the second image is regressed using the second bounding box, thereby achieving accurate matching of targets in the binocular image. It is not necessary to perform target detection on both images of the binocular image and then use a matching algorithm to match the targets detected in the two images. This greatly reduces the computational overhead of target matching in the binocular image, solves the problem of the complexity of the matching algorithm set by multiple features and thresholds in the prior art, avoids the limitations of target matching in the binocular image, and improves the efficiency, accuracy and robustness of target matching in the binocular image.

[0058] Example 2

[0059] Figure 2A This is a flowchart of a binocular image matching method provided in Embodiment 2 of the present invention. Figure 2B This is a schematic diagram illustrating the principle of the binocular image matching process provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, this embodiment mainly provides a detailed explanation of the specific regression process from the second bounding box to the third bounding box of the target within the second image.

[0060] Specifically, such as Figure 2A As shown, this embodiment may include the following steps:

[0061] S210, Perform target detection on the first image of the binocular image to obtain the first bounding box of the target in the first image.

[0062] S220, determine the second bounding box corresponding to the first bounding box in the second image of the stereo image.

[0063] S230, a pre-built feature extraction network is used to generate feature maps of the second image.

[0064] Optionally, when regressing the second bounding box of the target in the second image, it is necessary to analyze a certain mapping relationship between the second bounding box and the actual features of the target in the second image, so it is necessary to extract the target features in the second image.

[0065] In this embodiment, as Figure 2B As shown, a corresponding feature extraction network is pre-built, and the second image is input into this feature extraction network. Then, the feature extraction network uses its pre-trained network parameters to perform feature analysis on the input second image, thereby generating a feature map of the second image, so as to accurately analyze the features of each target in the second image later.

[0066] S240, extract the candidate features corresponding to the second bounding box.

[0067] After obtaining the feature map of the second image, the relationship between the target and the feature map is analyzed based on the target's second bounding box within the second image. Then, according to this relationship, candidate features corresponding to the second bounding box can be accurately extracted from the feature map as the target features within the second image.

[0068] S250 uses a pre-built regression network to process candidate features and obtain the third bounding box.

[0069] Optionally, in order to perform accurate regression on the second bounding box, such as Figure 2B As shown, this embodiment pre-constructs a regression network. The input of the regression network is the candidate features of the second bounding box, and the output is the deviation between the second bounding box and the third bounding box. The deviation may include the size deviation of the bounding box and the coordinate deviation of the same key point within the bounding box, and the key point may include the center point and / or diagonal vertex of the second and third bounding boxes.

[0070] For example, the deviation between the second and third bounding boxes in this embodiment may specifically include at least one of the following: the width ratio of the two bounding boxes (e.g., w3 / w2), the height ratio (e.g., h3 / h2), the ratio of the difference in the x-coordinates of the key points to the width of the second bounding box (e.g., (x3-x2) / w2), and the ratio of the difference in the y-coordinates of the key points to the height of the second bounding box (e.g., (y3-y2) / h2). Wherein, (x2, y2) is the coordinate of a key point in the second bounding box, (x3, y3) is the coordinate of the corresponding key point in the third bounding box, w2 and h2 are the width and height of the second bounding box, and w3 and h3 are the width and height of the third bounding box.

[0071] Of course, the bounding box can also be a three-dimensional box. In this case, the deviation can be the ratio of the length of the two bounding boxes, the ratio of the width of the two bounding boxes, the ratio of the height of the two bounding boxes, the ratio of the difference in the x-coordinate of the key points to the length of the second bounding box, the ratio of the difference in the y-coordinate of the key points to the width of the second bounding box, and the ratio of the difference in the y-coordinate of the key points to the height of the second bounding box. In addition, this disclosure can also take the logarithm (log) of the size-related ratios (i.e., the length ratio, the width ratio, and the height ratio) to reduce the impact of changes in object scale on the regression.

[0072] Therefore, after extracting the candidate features corresponding to the second bounding box, these candidate features can be input into the regression network. The network's pre-trained parameters then perform regression processing on these candidate features, outputting the deviation between the second and third bounding boxes. This deviation is then used to adjust the second bounding box accordingly, thus obtaining the third bounding box of the target within the second image.

[0073] The technical solution provided in this embodiment can directly output the deviation between the second bounding box and the third bounding box when performing regression processing on the candidate features of the second bounding box through a pre-constructed regression network. This makes it easy to quickly determine the various differences between the second and third bounding boxes using the deviation. When it is necessary to use the difference parameters between the second and third bounding boxes to calculate other parameters, they can be directly obtained from the deviation, thereby improving the efficiency of determining the difference parameters when matching targets in binocular images.

[0074] Example 3

[0075] Figure 3 This is a flowchart of a binocular image matching method provided in Embodiment 3 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, considering that the second bounding box can be any of the following candidate boxes: 1) a box at the same image position as the first bounding box, denoted as the first candidate box; 2) a box obtained by translating the first bounding box and estimating the disparity, denoted as the second candidate box; 3) a box obtained by binocularly correcting the first bounding box and projecting it back to the uncorrected second image coordinate system, denoted as the third candidate box; 4) a box obtained by binocularly correcting and translating the first bounding box and estimating the disparity, and projecting it back to the uncorrected second image coordinate system, denoted as the fourth candidate box. When obtaining the second and fourth candidate boxes, it is necessary to translate the first bounding box according to the estimated disparity. Therefore, before determining the second bounding box corresponding to the first bounding box in the second image of the binocular image, it is first necessary to determine the estimated disparity of the target within the binocular image. Accordingly, this embodiment mainly provides a detailed explanation of the specific calculation process of the estimated disparity of the target within the binocular image.

[0076] For details, please refer to Figure 3 The method may include the following steps:

[0077] S310, Perform target detection on the first image of the binocular image to obtain the first bounding box of the target in the first image and the target category.

[0078] Optionally, this embodiment can incorporate the principle of monocular ranging into binocular ranging to estimate the disparity of the binocular image. The monocular ranging formula can be: Where w is the pixel width of the target, f is the focal length, W is the actual width of the target, and depth is the target depth. The formula for binocular ranging is as mentioned in Example 2 above:

[0079] Then, combining the two formulas above, we can obtain: in This represents the translation coefficient of the first bounding box. Therefore, in this embodiment, the estimated disparity is related to the translation coefficient and the pixel width of the target. The pixel width of the target can be taken as the width of the detection box. Furthermore, when the target is a vehicle body, the pixel width can be taken as the width of the vehicle body, such as the width of the rear of the vehicle or the width of the cross-section of the vehicle body. In this case, a specific algorithm can be used to accurately extract the corresponding width of the target in the binocular image.

[0080] Moreover, the translation coefficient is related to the baseline of the stereo camera and the actual width of the target. The baseline of the stereo camera can be determined in advance, and considering that the actual width of each target in the same category is roughly the same, the translation coefficient can be an unknown quantity related to the target category. Different target categories will be assigned corresponding translation coefficients.

[0081] Therefore, in order to determine the corresponding translation coefficient, this embodiment will also detect the category of each target in the first image when performing target detection, so as to obtain the translation coefficient corresponding to the category of the target.

[0082] S320, obtain the translation coefficient corresponding to the category of the target.

[0083] According to the pre-defined correspondence between target categories and translation coefficients, after detecting the category of the target in the first image, this embodiment can obtain the translation coefficient corresponding to the target category, so as to calculate the estimated disparity of the target in the second image based on the translation coefficient and the pixel width of the target.

[0084] It should be noted that, regarding the correspondence between target categories and translation coefficients, this embodiment can use the pixel width and actual disparity of each target under each category in a large number of historical stereo images to analyze the relationship between target categories and translation coefficients. That is, the actual disparity and pixel width of the same target that has been marked and matched in historical stereo images are determined; based on the pixel width and actual disparity of each target under the same category that has been marked, the translation coefficient corresponding to that category is calculated.

[0085] Specifically, in historical stereo images, the first bounding box of each target within the first image and the third bounding box within the second image are accurately marked, thereby enabling accurate matching of the same target within the stereo images. Furthermore, the first and third bounding boxes marked in the historical stereo images are bounding boxes after image correction. Therefore, based on the first and third bounding boxes of the matched targets in the historical stereo images, the actual disparity and pixel width of each matched target can be obtained. Then, the targets in the historical stereo images are categorized to obtain the targets within each category. Those skilled in the art can set the classification granularity of the target major or minor categories as needed; this disclosure does not impose any limitations on this. Moreover, since there are multiple targets with actual disparities and pixel widths within each category, the historical translation coefficient of each target within each category is obtained. Then, by averaging the historical translation coefficients of each target within each category, the corresponding translation coefficient for that category can be obtained.

[0086] For example, this embodiment can use the following formula: To calculate the translation coefficient for each category. Where k i Let be the translation coefficient corresponding to the i-th category. Let i be the number of targets in the i-th category. The actual disparity of the j-th target in the i-th category. Let be the pixel width of the j-th target in the i-th category.

[0087] Furthermore, considering that historical stereo images may contain corresponding stereo camera baselines, and the current stereo image to be matched also has a corresponding stereo camera baseline, if different stereo cameras are used, the stereo camera baselines in the translation coefficients will also be different. Therefore, to ensure the accuracy of the translation coefficients, this embodiment, when calculating the translation coefficients corresponding to each category based on the pixel width and actual disparity of each target under the same category that has been marked, also refers to the stereo camera baselines marked in historical stereo images, dividing the translation coefficients in this embodiment into two parts: the stereo camera baseline and the sub-translation coefficients. According to the formula... It can be determined that the sub-translation coefficient in this embodiment is the actual width of the target under each category. Therefore, when calculating the sub-translation coefficient t corresponding to that category based on the pixel width and actual disparity of each target under the same labeled category, and the corresponding binocular camera baseline, the calculation formula referenced is:

[0088]

[0089] Among them, t i The sub-translation coefficient corresponding to the i-th category; W represents the baseline of the stereo camera corresponding to the j-th target in the i-th category;i The actual width of the i-th category is not statistically analyzed. This disclosure obtains the t-value (i.e., 1 / W value) of different target categories under each category through mean statistics, and the baseline of the stereo camera in actual application. The translation coefficient k of the stereo image to be matched can then be obtained through a linear relationship.

[0090] Accordingly, after calculating the sub-translation coefficient corresponding to each target category, the translation coefficient corresponding to the target category can be determined based on this sub-translation coefficient and the stereo camera baseline corresponding to the current stereo image to be matched. In other words, the formula... To calculate the translation coefficient corresponding to the category of each target in the second image.

[0091] S330 determines the estimated disparity of the target in the binocular image based on the translation coefficient and the target's pixel width.

[0092] The pixel width of the target is determined based on the width of the target's first bounding box. Then, the translation coefficient of the target within the second image and the pixel width of the target are substituted into the formula. In the process, the estimated disparity of the target in the binocular image is calculated to accurately translate the first bounding box in the second image, ensuring the accuracy of the second bounding box.

[0093] S340, determine the second bounding box corresponding to the first bounding box in the second image of the stereo image.

[0094] S350, using the second bounding box to regress the third bounding box of the target within the second image.

[0095] The technical solution provided in this embodiment, by performing target detection on the first image of the binocular image, can obtain the first bounding box of the target in the first image. At this time, the second bounding box corresponding to the first bounding box can be directly determined in the second image of the binocular image. Then, the third bounding box of the target in the second image is regressed using the second bounding box, thereby achieving accurate matching of targets in the binocular image. It is not necessary to perform target detection on both images of the binocular image and then use a matching algorithm to match the targets detected in the two images. This greatly reduces the computational overhead of target matching in the binocular image, solves the problem of the complexity of the matching algorithm set by multiple features and thresholds in the prior art, avoids the limitations of target matching in the binocular image, and improves the efficiency, accuracy and robustness of target matching in the binocular image.

[0096] Example 4

[0097] Figure 4 This is a schematic diagram of the structure of a binocular image matching device provided in Embodiment 4 of the present invention, as shown below. Figure 4 As shown, the device may include:

[0098] The target detection module 410 is used to perform target detection on the first image to obtain the first bounding box of the target in the first image;

[0099] The region determination module 420 is used to determine the second bounding box corresponding to the first bounding box in the second image of the stereo image;

[0100] The regression module 430 is used to regress the third bounding box of the target within the second image using the second bounding box.

[0101] The technical solution provided in this embodiment, by performing target detection on the first image of the binocular image, can obtain the first bounding box of the target in the first image. At this time, the second bounding box corresponding to the first bounding box can be directly determined in the second image of the binocular image. Then, the third bounding box of the target in the second image is regressed using the second bounding box, thereby achieving accurate matching of targets in the binocular image. It is not necessary to perform target detection on both images of the binocular image and then use a matching algorithm to match the targets detected in the two images. This greatly reduces the computational overhead of target matching in the binocular image, solves the problem of the complexity of the matching algorithm set by multiple features and thresholds in the prior art, avoids the limitations of target matching in the binocular image, and improves the efficiency, accuracy and robustness of target matching in the binocular image.

[0102] Furthermore, the second bounding box mentioned above can be any of the following candidate boxes:

[0103] The box at the same image position as the first bounding box is denoted as the first candidate box;

[0104] The box obtained by translating the first bounding box and estimating the disparity is denoted as the second candidate box;

[0105] The box obtained by binocular correction of the first bounding box and projection back to the uncorrected second image coordinate system is denoted as the third candidate box.

[0106] The first bounding box is corrected by binoculars and its parallax is estimated by translation. The resulting box is then projected back into the uncorrected second image coordinate system and denoted as the fourth candidate box.

[0107] Furthermore, when the baseline and focal length of the binocular camera corresponding to the binocular image to be matched meet the preset low parallax condition, the second bounding box is the first candidate box.

[0108] Furthermore, the regression module 430 mentioned above can be specifically used for:

[0109] A pre-constructed feature extraction network is used to generate a feature map of the second image;

[0110] Extract candidate features corresponding to the second bounding box;

[0111] The candidate features are processed using a pre-constructed regression network to obtain the third bounding box.

[0112] Furthermore, the input to the regression network is the candidate features of the second bounding box, and the output is the deviation between the second bounding box and the third bounding box; the deviation includes the size deviation and the coordinate deviation of the key points, and the key points include the center point and / or the diagonal vertex.

[0113] Furthermore, the aforementioned deviation may include at least one of the following:

[0114] The ratios of the widths and heights of the two bounding boxes, the ratio of the difference in the x-coordinates of the key points to the width of the second bounding box, and the ratio of the difference in the y-coordinates of the key points to the height of the second bounding box.

[0115] Furthermore, the target detection result may also include the target category, and the binocular image matching device may also include: a disparity estimation calculation module;

[0116] The estimated disparity calculation module is used to obtain the translation coefficient corresponding to the category of the target; and to determine the estimated disparity of the target in the binocular image based on the translation coefficient and the pixel width of the target.

[0117] The pixel width of the target is determined based on the width of the target's first bounding box.

[0118] Furthermore, the aforementioned binocular image matching device may also include:

[0119] The historical parameter determination module is used to determine the actual disparity and pixel width of the same target that has been marked and matched in historical binocular images;

[0120] The translation coefficient calculation module is used to calculate the translation coefficient corresponding to each category based on the pixel width and actual disparity of each target under the same labeled category.

[0121] Furthermore, the aforementioned historical binocular images can also be marked with corresponding binocular camera baselines. The aforementioned translation coefficient calculation module can be specifically used for:

[0122] Calculate the sub-translation coefficient corresponding to the category based on the pixel width and actual disparity of each target under the same labeled category, as well as the corresponding binocular camera baseline;

[0123] Accordingly, the aforementioned parallax estimation module can be specifically used for:

[0124] Based on the sub-translation coefficients and the stereo camera baseline corresponding to the current stereo image to be matched, the translation coefficients corresponding to the category of the target are determined.

[0125] Furthermore, the aforementioned binocular image matching device may also include:

[0126] The actual disparity calculation module is used to calculate the actual disparity of the target in the binocular image based on the first bounding box and the third bounding box;

[0127] The target depth calculation module is used to calculate the depth information of the target based on the actual parallax, the binocular camera baseline and focal length corresponding to the binocular image.

[0128] The binocular image matching device provided in this embodiment can be applied to the binocular image matching method provided in any of the above embodiments, and has the corresponding functions and beneficial effects.

[0129] Example 5

[0130] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present invention. Figure 5 As shown, the electronic device includes a processor 50, a storage device 51, and a communication device 52; the number of processors 50 in the electronic device can be one or more. Figure 5 Taking a processor 50 as an example; the processor 50, storage device 51, and communication device 52 of the electronic device can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0131] Storage device 51, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules. Processor 50 executes various functional applications and data processing of electronic devices by running the software programs, instructions, and modules stored in storage device 51, thereby realizing the aforementioned binocular image matching method.

[0132] Storage device 51 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on terminal usage. Furthermore, storage device 51 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, storage device 51 may further include memory remotely configured relative to the multifunction controller 50, which can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0133] The communication device 52 can be used to realize network connection or mobile data connection between devices.

[0134] The electronic device provided in this embodiment can be used to execute the binocular image matching method provided in any of the above embodiments, and has the corresponding functions and beneficial effects.

[0135] Example 6

[0136] Embodiment Six of the present invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, can implement the binocular image matching method in any of the above embodiments. The method may include:

[0137] Target detection is performed on the first image of the binocular image to obtain the first bounding box of the target in the first image;

[0138] In the second image of the binocular image, determine the second bounding box corresponding to the first bounding box;

[0139] The third bounding box of the target within the second image is regressed using the second bounding box.

[0140] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the method operations described above, but can also perform related operations in the binocular image matching method provided in any embodiment of the present invention.

[0141] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0142] It is worth noting that in the embodiments of the binocular image matching device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0143] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A binocular image matching method, characterized in that, The method includes: Target detection is performed on the first image of the binocular image to obtain the first bounding box of the target in the first image and the category of the target; In the second image of the binocular image, determine the second bounding box corresponding to the first bounding box; The third bounding box of the target within the second image is regressed using the second bounding box; The method further includes: Obtain the translation coefficient corresponding to the category of the target; Based on the translation coefficient and the pixel width of the target, the estimated disparity of the target in the binocular image is determined, wherein the pixel width of the target is determined according to the width of the target's first bounding box.

2. The method according to claim 1, characterized in that, The second bounding box is any one of the following candidate boxes: The box at the same image position as the first bounding box is denoted as the first candidate box; The box obtained by translating the first bounding box and estimating the disparity is denoted as the second candidate box; The box obtained by binocular correction of the first bounding box and projection back to the uncorrected second image coordinate system is denoted as the third candidate box. The first bounding box is corrected by binoculars and its parallax is estimated by translation. The resulting box is then projected back into the uncorrected second image coordinate system and denoted as the fourth candidate box.

3. The method according to claim 2, characterized in that, When the baseline and focal length of the stereo camera corresponding to the stereo image to be matched meet the preset low parallax condition, the second bounding box is the first candidate box.

4. The method according to claim 1, characterized in that, Using the second bounding box, regress the third bounding box of the target within the second image, including: A pre-constructed feature extraction network is used to generate a feature map of the second image; Extract candidate features corresponding to the second bounding box; The candidate features are processed using a pre-constructed regression network to obtain the third bounding box.

5. The method according to claim 4, characterized in that, The input to the regression network is the candidate features of the second bounding box, and the output is the deviation between the second bounding box and the third bounding box; the deviation includes the size deviation and the coordinate deviation of the key points, and the key points include the center point and / or the diagonal vertex.

6. The method according to claim 5, characterized in that, The deviation includes at least one of the following: The ratios of the widths and heights of the two bounding boxes, the ratio of the difference in the x-coordinates of the key points to the width of the second bounding box, and the ratio of the difference in the y-coordinates of the key points to the height of the second bounding box.

7. The method according to claim 6, characterized in that, Before obtaining the translation coefficient corresponding to the category of the target, the method further includes: Determine the actual disparity and pixel width of the same target that has been labeled and matched in historical binocular images; Calculate the translation coefficient corresponding to each category based on the pixel width and actual disparity of each target in the same labeled category.

8. The method according to claim 7, characterized in that, The historical binocular images also include corresponding binocular camera baselines. Based on the pixel width and actual disparity of each target within the labeled category, the translation coefficient corresponding to that category is calculated, including: Calculate the sub-translation coefficient corresponding to the category based on the pixel width and actual disparity of each target under the same labeled category, as well as the corresponding binocular camera baseline; Accordingly, obtaining the translation coefficient corresponding to the category of the target includes: Based on the sub-translation coefficients and the stereo camera baseline corresponding to the current stereo image to be matched, the translation coefficients corresponding to the category of the target are determined.

9. The method according to claim 1, characterized in that, After regressing the third bounding box of the target within the second image using the second bounding box, the method further includes: Calculate the actual disparity of the target in the binocular image based on the first bounding box and the third bounding box; The depth information of the target is calculated based on the actual parallax, the binocular camera baseline and focal length corresponding to the binocular image.

10. A binocular image matching device, characterized in that, include: The target detection module is used to perform target detection on the first image to obtain the first bounding box of the target in the first image and the category of the target; The region determination module is used to determine the second bounding box corresponding to the first bounding box in the second image of the stereo image; The regression module is used to regress the third bounding box of the target within the second image using the second bounding box; The binocular image matching device further includes a disparity estimation calculation module, used for: Obtain the translation coefficient corresponding to the category of the target; Based on the translation coefficient and the pixel width of the target, the estimated disparity of the target in the binocular image is determined, wherein the pixel width of the target is determined according to the width of the target's first bounding box.

11. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the binocular image matching method as described in any one of claims 1-9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the binocular image matching method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Method and system for disparity adjustment during stereoscopic zoom

    CN103270760A

  • Apparatus and method for processing image pair obtained from stereo camera

    US20180041747A1