Method, device and system for obtaining binocular disparity map for small target

Through technical means such as hierarchical feature extraction, attention weight screening, and adaptive multimodal cross-entropy loss function, the problem of low parallax matching accuracy of small targets is solved, the accuracy of the generation of parallax maps is improved, and the effect of three-dimensional reconstruction and autonomous driving is improved.

CN120259399AActive Publication Date: 2025-07-04BEIJING SMARTER EYE TECH CO LTD

Patent Information

Application Number
CN202510653994.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-04
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The prior art generates a parallax map, especially for small targets, with low parallax matching accuracy, which affects the accuracy and reliability of three-dimensional reconstruction and autonomous driving.

Method used

The hierarchical feature extraction module, attention weight screening, convolutional context information fusion and adaptive multimodal cross-entropy loss function are used to improve the parallax matching accuracy of small targets through feature map compression, cost volume construction and disparity estimation.

Benefits of technology

It significantly improves the parallax matching accuracy of small targets, improves the accuracy of three-dimensional reconstruction and the safety of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259399A_ABST
    Figure CN120259399A_ABST
Patent Text Reader

Abstract

The invention discloses a binocular disparity map obtaining method, device and system for a small target, and is used for improving the disparity optimization effect of the small target, and the method comprises the steps: carrying out the hierarchical feature extraction of a left view and a right view, obtaining a feature map, compressing the feature map to 32 channels through the convolution operation integrated with the structure information, and obtaining a binocular disparity map; therefore, an initial splicing body is constructed; obtaining attention weights of the stereo images corresponding to the left view and the right view, screening the initial splicing bodies by using the attention weights, and constructing a cost body according to a screening result and initial parallax loss; fusing the convolutional context information with the preliminarily aggregated intermediate features to perform cost aggregation on the cost body to obtain an aggregation result; and adopting a self-adaptive multi-modal cross entropy loss function to supervise an aggregation result, and performing multi-modal output on the supervised aggregation result through a multi-modal parallax estimator to obtain a binocular parallax graph of the left view and the right view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of assisted driving technology, and in particular, to a method, device, and system for obtaining binocular disparity maps for small targets. Background Art

[0002] Stereo matching technology has a wide range of applications in the field of computer vision, such as 3D reconstruction, autonomous driving, etc. In the field of autonomous driving, stereo matching technology obtains the depth information of the surrounding environment through binocular cameras, realizes accurate obstacle detection and distance estimation, and provides important support for vehicle navigation and obstacle avoidance. In addition, stereo matching technology is also combined with the road preview system to improve the driving stability and comfort of the vehicle. The road preview system uses binocular stereo vision technology to scan and identify the road conditions in front of the vehicle in real time, and obtains information such as the flatness of the road surface, obstacles, and the slope of the road. This information is transmitted to the vehicle's electronic control unit (ECU), and the ECU adjusts the damping characteristics of the suspension system according to this information to adapt to the road conditions ahead. For example, if a speed bump is previewed ahead, the ECU will adjust the suspension system in advance to make it stiffer to reduce the impact when the vehicle passes through. This technology can significantly improve the driving quality of the vehicle and the comfort of the passengers.

[0003] However, there are some limitations in the existing technology when generating disparity maps. Due to the limitations of the network structure or algorithm, geometric structure information is easily lost during the disparity generation process, especially at the edge details of the image, and the disparity matching accuracy of small targets is relatively low. This information loss will cause deviations in the details of the 3D reconstruction model, affecting the overall accuracy and reliability. For example, in autonomous driving, if the disparity matching of small targets (such as road signs, speed bumps, etc.) is inaccurate, it may lead to incorrect distance estimation of these targets by the vehicle, thus affecting driving safety.

[0004] In view of this, this application is proposed. Summary of the Invention

[0005] The main purpose of this application is to disclose a method, device, and system for obtaining binocular disparity maps for small targets, which is used to solve the problem of relatively low disparity matching accuracy of small targets in the existing technology.

[0006] To achieve the above object, according to one aspect of this application, a method for obtaining a binocular disparity map for small targets is provided, and the following technical solutions are adopted:

[0007] A binocular disparity map acquisition method for small targets includes: performing hierarchical feature extraction on the left and right views to obtain feature maps, compressing the feature maps to 32 channels through a feature extraction module incorporating structural information, thereby constructing an initial mosaic; obtaining the attention weights of the corresponding stereo images of the left and right views, using the attention weights to screen the initial mosaic, and constructing a cost volume according to the screening results and the initial disparity loss; fusing convolutional context information with the intermediate features obtained after preliminary aggregation to perform cost aggregation on the cost volume to obtain an aggregation result; using an adaptive multi-modal cross-entropy loss function to supervise the aggregation result, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

[0008] Further, the performing hierarchical feature extraction on the left and right views to obtain feature maps, compressing the feature maps to 32 channels through a feature extraction module incorporating structural information, thereby constructing an initial mosaic includes: performing downsampling on the input left and right views respectively with a convolutional kernel of a preset size and corresponding strides to reduce the image size; using 16 residual layers to extract unary features at a resolution of 1 / 4, that is layer feature maps, layer feature maps and layer feature maps each contain 3 residual layers, expanding the receptive field and obtaining semantic information by increasing the number of channels; on the basis of the original convolutional feature layer, local structural information LCS is incorporated into each layer, and all feature maps at a resolution of 1 / 4 are fused to generate a feature map with 320 channels, and by compressing the feature map to 32 channels, an initial mosaic is constructed.

[0009] Further, the using 16 residual layers to extract unary features at a resolution of 1 / 4, that is layer feature maps, layer feature maps and layer feature maps each contain 3 residual layers, expanding the receptive field and obtaining semantic information by increasing the number of channels includes: at each convolutional feature extraction layer, calculating the structural information LCS of different dilation rates, increasing the structural information LCS of feature extraction, fusing multiple local structural feature layers with the original convolutional feature layer, and the fused features will learn by themselves to balance appearance information and structural information; in the structural information LCS, each channel represents the relationship between a central pixel and its adjacent pixels, and its formula is:

[0010] where represents the pixel point on the k-th channel of the LCS, is the original convolutional feature, represents the offset of the k-th neighboring point, is a function used to measure the cosine similarity between two vectors. In a convolutional neural network, low-level features contain rich texture information, while high-level features usually reveal semantic clues. The neighborhood relationship is calculated through the features extracted from different layers.

[0011] Further, the obtaining of the attention weights for the corresponding stereo images of the left and right views, using the attention weights to screen the initial spliced body, and constructing the cost volume according to the screening result and the initial disparity loss includes: Step 1: Construction of the initial cost volume; The size of the stereo image pair input to the network is H×W×3. After each image goes through the feature extraction step, the unary feature maps of the left view and the right view are obtained respectively. and ; The feature map ( ) has a size of ×H / 4×W / 4, where = 32, representing the number of channels, H: height, W: width. The initial cost volume is formed by splicing the feature maps and at each disparity level; Step 2: Generation of attention weights; Three different levels of feature maps are obtained from the feature extraction module, with the number of channels being 64, 128, and 128 respectively; For each pixel at a specific level, a dilation block with a predefined size and adaptive learning weights is used to calculate the matching cost; By controlling the dilation rate, it is ensured that the range of the window is related to the feature map level, and at the same time, to keep the same number of pixels in calculating the similarity of the central pixel, the similarity of two corresponding pixels is the weighted sum of the correlations of the corresponding pixels within the window; Step 3: Divide the features into groups and calculate the correlation map for each group; The channels are evenly divided into groups, where the first 8 groups come from the layer feature map, the middle 16 groups come from the layer feature map, and the last 16 groups come from the layer feature map; The feature maps at different levels do not interfere with each other. The g-th group of features is represented as , . The matching cost volume of the g-th group, the matching weight of the g-th group of features, is adaptively learned during the training process. The final multi-level block matching cost volume is obtained by connecting the matching costs of all levels; The obtained multi-level block matching cost volume is normalized by applying two 3D convolutions and a 3D hourglass network, and then another convolutional layer is used to compress the channels to 1, thereby obtaining the attention weights.

[0012] Further, fusing the intermediate features obtained after convolving the context information and preliminary aggregation to perform cost aggregation on the cost volume, and the obtained aggregation result includes: given the left view, the context features obtained in the convolutional feature extraction step , and the geometric features obtained by preliminary aggregation of the attention cost volume ; expanding the size of to , denoted as , where N represents the number of channels of the feature maps at different levels, and D represents the maximum disparity; generating spatial attention weights by exploring the spatial relationship between the convolutional context and geometric features, using the attention mechanism to fuse features, so as to adaptively select the important regions of the features; using an aggregation strategy based on the hourglass structure to process the attention cost volume, and realizing multi-scale fusion of features through step-by-step downsampling and upsampling.

[0013] Further, using the adaptive multi-modal cross-entropy loss function to supervise the aggregation result, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator, and the obtained binocular disparity maps of the left and right views include: the training process of the network includes: Phase 1: Only train the attention weight part, and perform supervision using the modeling based on the ground truth probability; Phase 2: The attention weight part is frozen to focus on training other parts of the network; Phase 3: The entire network is trained together to achieve the improvement of the overall performance; in each phase, the adaptive multi-modal cross-entropy loss function is used to supervise the aggregation result, and the cross-entropy loss of this adaptive multi-modal is expressed as:

[0014] where is the predicted disparity probability distribution obtained through the attention weight A, represents coefficient, is the i-th predicted disparity probability distribution output during the training process, represents coefficient, represents the above-mentioned ground truth probability modeling.

[0015] Further, the method of using the adaptive multi-modal cross-entropy loss function to supervise the aggregation result and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views further includes: separating each modality in the multi-modal output and calculating the cumulative probability respectively; each modality represents a potential matching object with a specific depth, and its cumulative probability represents the possibility of the object being matched; adopting an object-level WTA strategy to select the modality with the highest cumulative probability as the dominant modality and performing normalization processing on this modality: finally, estimating the disparity using weighted average operation, estimating the multi-modal output of the network through a multi-modal disparity estimator, and obtaining the final disparity map.

[0016] According to another aspect of the present application, a binocular disparity map acquisition device for small targets is provided, and the following technical solutions are adopted:

[0017] A binocular disparity map acquisition device for small targets includes: an extraction module, configured to perform hierarchical feature extraction on the left and right views to obtain feature maps, and compress the feature maps to 32 channels through a feature extraction module incorporating structural information, thereby constructing an initial mosaic; a construction module, configured to obtain the attention weights of the corresponding stereo images of the left and right views, use the attention weights to screen the initial mosaic, and construct a cost volume according to the screening result and the initial disparity loss; an aggregation module, configured to perform cost aggregation on the cost volume by fusing convolutional context information and intermediate features obtained after preliminary aggregation to obtain an aggregation result; a supervision module, configured to use an adaptive multi-modal cross-entropy loss function to supervise the aggregation result, and perform multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

[0018] Further, the extraction module includes: a sampling module, configured to perform downsampling on the input left and right views respectively with a convolution kernel of a preset size and corresponding strides to reduce the image size; an acquisition module, configured to extract unary features at a resolution of 1 / 4 using 16 residual layers, layer feature maps, layer feature maps and layer feature maps each include 3 residual layers, expand the receptive field and obtain semantic information by increasing the number of channels; a fusion module, configured to fuse all feature maps at a resolution of 1 / 4 to generate a feature map with 320 channels, and construct an initial mosaic by compressing the feature map to 32 channels.

[0019] According to yet another aspect of the present application, a binocular disparity map acquisition system for small targets is provided, and the following technical solutions are adopted:

[0020] A binocular disparity map acquisition system for small targets includes the above-mentioned device.

[0021] In this application, by adding local structure information (LCS) to the feature extraction module, introducing context information (CCF) to the cost aggregation module, and using an adaptive multi-modal cross-entropy loss function (AML) and a multi-modal disparity estimator for optimization (MDE), this application can effectively improve the disparity matching accuracy of small targets, thereby improving the accuracy of 3D reconstruction and the safety of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a flowchart of a method for obtaining a binocular disparity map for small targets according to an embodiment of the present application;

[0024] Figure 2 It is a flowchart of another method for obtaining a binocular disparity map for small targets according to an embodiment of the present application;

[0025] Figure 3 It is a structural diagram of a device for obtaining a binocular disparity map for small targets according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following will describe the embodiments of the present application in detail with reference to the accompanying drawings. However, the present application can be implemented in many different ways defined and covered by the claims.

[0027] Figure 1 It is a flowchart of a method for obtaining a binocular disparity map for small targets according to an embodiment of the present application.

[0028] See Figure 1 As shown, a method for obtaining a binocular disparity map for small targets includes:

[0029] S101: Perform hierarchical feature extraction on the left and right views to obtain a feature map, and compress the feature map to 32 channels through a feature extraction module incorporating local structure information, thereby constructing an initial mosaic;

[0030] S103: Obtain the attention weights of the corresponding stereo images of the left and right views, use the attention weights to screen the initial mosaic, and construct a cost volume according to the screening results and the initial disparity loss;

[0031] S105: Fuse the convolutional context information with the intermediate features obtained after preliminary aggregation to perform cost aggregation on the cost volume, and obtain an aggregation result;

[0032] S107: Use an adaptive multi-modal cross-entropy loss function to supervise the aggregation result, and perform multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

[0033] First, the binocular disparity map acquisition method for small targets proposed in this application includes five parts: feature extraction, cost volume construction, cost aggregation, loss function, and disparity estimation.

[0034] In step S101, hierarchical feature extraction is performed on the left and right views to obtain feature maps, and the feature maps are compressed to 32 channels through a feature extraction module incorporating local structure information, thereby constructing an initial mosaic. Specifically, in the convolutional feature extraction part, a three-stage ResNet architecture is adopted, and the features are divided into three levels. First, three 3×3 convolutional kernels are used to downsample the input image at strides of 2, 1, and 1 respectively to reduce the image size. Next, 16 residual layers are used to extract unary features at a resolution of 1 / 4, that is, layer feature maps, and layer feature maps, each containing 3 residual layers, to expand the receptive field and obtain richer semantic information by increasing the number of channels. Finally, the feature maps of all layers at a resolution of 1 / 4 ( 、 、 ) are fused to generate a feature map with 320 channels ( = 320) for calculating attention weights. Subsequently, the feature map is compressed to 32 channels through a convolutional operation, thereby constructing an initial mosaic.

[0035] In step S103, obtain the attention weights of the corresponding stereo images of the left and right views, use the attention weights to screen the initial mosaic, and construct the cost volume according to the screening result and the initial disparity loss. Specifically, first is the construction of the initial cost volume. The size of the stereo image pair input to the network is H×W×3. After each image passes through the feature extraction step, the unary feature maps of the left and right images and are obtained respectively. The feature map ( ) has a size of × H / 4 × W / 4 (where = 32, representing the number of channels, H: height, W: width). The initial cost volume is obtained by, at each disparity level, and It is formed by splicing. Then, the attention weights are generated to screen the initial spliced body to highlight useful information and suppress irrelevant information. For this purpose, the geometric information in the correlation between stereo image pairs is extracted through multi-level adaptive block matching, thereby generating attention weights. Three different levels of feature maps are obtained from the feature extraction module , and , and their numbers of channels are 64, 128, and 128 respectively. For each pixel at a specific level, a dilated block with a predefined size and adaptive learning weights is used to calculate the matching cost. By controlling the dilation rate, it is ensured that the range of the window is related to the feature map level. At the same time, to keep the same number of pixels in calculating the similarity of the central pixel, the similarity of two corresponding pixels is the weighted sum of the correlations of the corresponding pixels within the window. Then, adopting the grouping idea of GwcNet, the features are divided into groups, and the correlation maps are calculated group by group. The channels are evenly divided into groups ( = 40), where the first 8 groups come from feature map, the middle 16 groups come from feature map, and the last 16 groups come from feature map. Feature maps at different levels do not interfere with each other. Represent the g-th group of features as , , and the matching cost volume of the g-th group is expressed as:

[0036]

[0037] where, is the inner product, and d represents different disparity levels. Subsequently, calculate the matching cost volumes of different levels k = 1, 2, 3) as:

[0038]

[0039] where, is a nine-point coordinate set that defines the window range on the k-level feature map, represents the matching weight of the k-th feature layer and the g-th group of features, and is adaptively learned during the training process. The final multi-level block matching cost volume is obtained by concatenating the matching costs of all levels:

[0040]

[0041] The obtained multi-level block matching cost volume is normalized by applying two 3D convolutions and a 3D hourglass network, and then another convolutional layer is used to compress the channels to 1 to obtain the attention weights. To obtain accurate attention weights for different disparities to filter the initial stitching cost volume, the ground truth disparity is used to supervise A. Specifically, the probability distribution obtained from A is normalized by the softmax function, and the adaptive multi-modal cross-entropy loss between the probability distribution and the ground truth disparity probability distribution is calculated to guide the network learning process, thereby obtaining accurate attention weights A. Finally, after obtaining the attention weights A, it is used to filter the redundant information in the initial stitching volume, thereby enhancing its representation ability. The final attention stitching volume The calculation formula on channel i is as follows:

[0042]

[0043] where ⊙ represents the Hadamard product.

[0044] In step S105, the cost volume is cost-aggregated by fusing the convolutional context information with the intermediate features obtained after preliminary aggregation, and the aggregation result is obtained. Specifically, in order to decode accurate and high-resolution feature information from low-resolution feature information with the assistance of context information, the present application proposes to fuse the convolutional context information and the intermediate features after preliminary aggregation (CCF) to achieve efficient and flexible cost aggregation. Given the context features obtained in the convolutional feature extraction step from the left-eye image , and the geometric features obtained after preliminary aggregation of the attention cost volume , expand the size of to , denoted as . Among them, N represents the number of channels of different-level feature maps, and D represents the maximum disparity value. The spatial attention weights are generated by exploring the spatial relationship between the convolutional context and the intermediate features, and the attention mechanism is used to fuse the features, thereby adaptively selecting the important regions of the features. The formula is as follows:

[0045]

[0046] An aggregation strategy based on the hourglass structure is used to process the attention cost volume, and multi-scale fusion of features is achieved by gradually downsampling and upsampling. First, the input cost volume Downsampling is performed through a series of 3D convolutional layers to gradually extract features at different scales and enhance the feature representation ability. After downsampling, CCF and an upsampling module are alternately used to decode high-resolution features. CCF adaptively fuses multi-scale features, introduces context information in the cost aggregation stage, enables the cost volume to better reflect the scene, further enhances features, and restores edges and weak texture regions.

[0047] In step S107, an adaptive multi-modal cross-entropy loss function is used to supervise the aggregation result, and a multi-modal disparity estimator is used to perform multi-modal output on the supervised aggregation result to obtain the binocular disparity maps of the left and right views. Specifically, in stereo matching, the smooth L1 loss is usually used to indirectly supervise the cost volume. However, this method limits the final performance improvement. By transforming the stereo matching problem into a classification task, the cross-entropy loss can directly supervise the probability distribution volume. Considering the uncertainty of disparity estimation and the multi-modal situation of the pixel probability distribution, this application adopts an adaptive multi-modal cross-entropy loss (AML), which provides more direct and effective supervision for the network, thus significantly improving the matching accuracy. The probability distribution of edge pixels should consist of multiple modes, and each mode represents a specific depth or disparity. Therefore, this application uses an adaptive multi-modal ground truth modeling method to generate an independent Laplace distribution for each potential depth of edge pixels, and then fuses these distributions to form a Laplace mixture model. The neighborhood of each pixel is used to complete this task. For each pixel with a labeled ground truth disparity, consider a local window of m×n (1×9) centered on this pixel. Then, the DBSCAN clustering algorithm is used to divide all the disparity values in the window into K non-overlapping subsets, and each subset corresponds to a different potential depth. The formula is as follows:

[0048]

[0049]

[0050] The Laplace distribution is discretized and normalized over the disparity candidates d∈{0,1,…,D - 1}, and 、 and are the mean, scale, and weight parameters of the k-th Laplace distribution respectively. Among them, is set to the average of the disparities within the cluster . Define to contain the central pixel to be modeled, and is replaced with the ground truth value of the central pixel to ensure the accuracy of supervision.

[0051]

[0052] Weight It is used to adjust the relative proportions of the obtained multiple modalities and can be distributed according to the local structure within the window. It is the fixed weight of the central pixel. Set , to ensure the dominance of the ground truth modality. The remaining weights are evenly distributed to the remaining 8 adjacent pixels. Taking as the index of the local structure, for example, a smaller corresponds to a finer structure and should thus have a smaller weight accordingly. For datasets with sparse ground truth like KITTI, only the valid disparities within the local window are calculated, and the equation becomes:

[0053]

[0054] For non-edge pixels with only one cluster within the window, is equal to 1, and the equation degenerates to a unimodal Laplace distribution.

[0055] Therefore, the final loss of the network uses the cross-entropy loss of adaptive multi-modalities, expressed as:

[0056]

[0057] Among them, is the predicted disparity probability distribution obtained through the attention weight A, represents 's coefficient, is the i-th predicted disparity probability distribution output during the training process, represents 's coefficient, represents the above-mentioned ground truth probability modeling.

[0058] The training process of the network is divided into three stages to ensure the effective learning and optimization of each part. In the first stage, only the attention weight part is trained, and supervised using the modeling based on the ground truth probability. In the second stage, the attention weight part is frozen to focus on training other parts of the network. Finally, in the third stage, the entire network is trained together to achieve the improvement of the overall performance. In each stage, the adaptive multi-modal cross-entropy loss function is adopted to ensure the stability and effectiveness of the training process.

[0059] Stereo matching networks trained with cross - entropy loss usually produce more multi - modal outputs than L1 loss, so the soft argmin function cannot be directly used to estimate the disparity. This application uses a multi - modal disparity estimator (MDE) to separate each modality in the multi - modal output and calculate their cumulative probabilities separately. Each modality represents a potential matching object with a specific depth, and its cumulative probability represents the likelihood of that object being a match. Therefore, an object - level WTA strategy is adopted to select the modality with the highest cumulative probability as the dominant modality. Subsequently, normalization is performed on this modality:

[0060]

[0061] Finally, the disparity is estimated using weighted average operations:

[0062]

[0063] By using MDE to estimate the multi - modal output of the network, the final disparity map can be obtained.

[0064] Through a series of innovative technologies, this application significantly improves the performance of the stereo matching algorithm in terms of the disparity matching accuracy and real - time performance of small targets. This application can more accurately capture the geometric features of small targets in images, such as edges, contours, etc. By optimizing the disparity matching process, this application reduces the loss of geometric structure information, making the generated disparity map more accurate in details. This provides higher - quality data input for 3D reconstruction, effectively reducing the deviation in details of the reconstructed 3D model and significantly improving the overall accuracy and reliability.

[0065] Figure 2 It is a flowchart of another method for obtaining a binocular disparity map for small targets described in the embodiments of this application.

[0066] See Figure 2 As shown, a method for obtaining a binocular disparity map for small targets includes:

[0067] Step 21: Input the left and right images;

[0068] Step 22: Weight - sharing feature extraction;

[0069] Step 221: Local structure feature extraction module;

[0070] Step 23: Cost volume construction;

[0071] Step 231: Initial disparity loss;

[0072] Step 232: Context information fusion module;

[0073] Step 24: Cost aggregation;

[0074] Step 25: Predict the disparity probability distribution;

[0075] Step 251: Adaptive multi-modal loss;

[0076] Step 26: Output the disparity map.

[0077] As a preferred embodiment, in the structural feature extraction part in Step 22, the structural information (LCS) composed of local cosine similarity is introduced. In LCS, each channel is used to represent the relationship between a central pixel and its adjacent pixels. The formula is:

[0078]

[0079] Where, represents the pixel point on the k-th channel of LCS, is the original convolutional feature, represents the offset of the k-th neighboring point, is a function used to measure the cosine similarity between two vectors. In this application, cosine similarity is selected to measure the relationship between two vectors because it has good numerical stability.

[0080] Considering that a smaller neighborhood helps in matching at disparity discontinuities, while a larger neighborhood helps in understanding the structure of textureless regions. To fully identify objects of different sizes, a square window (3×3) is designed to construct neighborhood relationships in different ranges. Since the pooling strategy may lead to loss of details, an expansion strategy is adopted to expand the receptive field. In a convolutional neural network, low-level features contain rich texture information, while high-level features usually reveal semantic clues. The neighborhood relationships calculated from features extracted from different layers are diverse. Therefore, in this application, at each convolutional feature extraction layer, the LCS with different dilation rates is calculated to increase the structural information of feature extraction. Multiple local structural feature layers are fused with the original convolutional feature layer, and the fused features will learn by themselves to balance appearance information and structural information.

[0081] Specifically, based on ACVNet, this application proposes a stereo matching algorithm that is both real-time and highly accurate. Compared with the original network, this application has achieved the following innovations:

[0082] On the basis of extracting appearance information from traditional convolutional features, the extraction of geometric structural features is introduced, which more accurately reveals the relationship between local pixels.

[0083] In the hourglass structure of cost aggregation, rich context information is incorporated, effectively retaining detailed features and improving the accuracy of matching.

[0084] In the three training stages of the network, an adaptive multi-modal cross-entropy loss is adopted, which can effectively guide the model to learn clear pixel distribution patterns, provide direct and efficient supervision for the network, and significantly improve the matching accuracy.

[0085] When performing disparity estimation and obtaining a disparity map, a multi-modal disparity estimation method is used, which can obtain more accurate results on the multi-modal output of the network.

[0086] This application applies an adaptive feature enhancement mechanism. In the feature extraction and cost aggregation stages, the parameters are dynamically adjusted according to the feature distribution of the input image and the requirements of the matching task.

[0087] Through the fusion-designed module, it significantly outperforms the baseline model ACVNet in the comparative experiments on the Scene Flow and KITTI datasets.

[0088] Figure 3 It is a structural diagram of a binocular disparity map acquisition device for small targets according to an embodiment of this application.

[0089] See Figure 3 As shown, a binocular disparity map acquisition device for small targets includes: an extraction module 30, configured to perform hierarchical feature extraction on the left and right views to obtain a feature map, and compress the feature map to 32 channels through a feature extraction module incorporating local structure information, thereby constructing an initial mosaic; a construction module 32, configured to obtain the attention weights of the corresponding stereo images of the left and right views, use the attention weights to screen the initial mosaic, and construct a cost volume according to the screening result and the initial disparity loss; an aggregation module 34, configured to fuse the convolutional context information and the intermediate features obtained after preliminary aggregation to perform cost aggregation on the cost volume to obtain an aggregation result; a supervision module 36, configured to supervise the aggregation result by using an adaptive multi-modal cross-entropy loss function, and perform multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

[0090] Optionally, the extraction module 30 includes: a sampling module (not shown in the figure), configured to perform downsampling on the input left and right views respectively with a convolutional kernel of a preset size and corresponding strides to reduce the image size; an acquisition module (not shown in the figure), configured to extract unary features at a resolution of 1 / 4 by using 16 residual layers, layer feature map, layer feature map and The layer feature maps each contain 3 residual layers, which expand the receptive field and obtain semantic information by increasing the number of channels, and each layer of feature maps incorporates local structural information; a fusion module (not shown in the figure) is used to fuse all the feature maps at 1 / 4 resolution to generate a feature map with 320 channels, and the feature map is compressed to 32 channels through a convolution operation, thereby constructing an initial mosaic.

[0091] A binocular disparity map acquisition system for small targets provided by the present application includes the above-mentioned device.

[0092] By adding local structural information (LCS) to the feature extraction module, introducing context information (CCF) to the cost aggregation module, and optimizing with an adaptive multi-modal cross-entropy loss function (AML) and a multi-modal disparity estimator (MDE), the present application can effectively improve the disparity matching accuracy of small targets, thereby improving the accuracy of 3D reconstruction and the safety of autonomous driving.

[0093] Only some exemplary embodiments of the present embodiment have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.

Claims

1. A binocular disparity map acquisition method for small targets, characterized in that Including: Performing hierarchical feature extraction on the left and right views to obtain a feature map, and compressing the feature map to 32 channels through a convolution operation incorporating structural information, thereby constructing an initial splicing body; Obtaining the attention weights of the corresponding stereo images of the left and right views, using the attention weights to screen the initial splicing body, and constructing a cost volume according to the screening result and the initial disparity loss; Fusing the convolutional context information with the preliminarily aggregated intermediate features to perform cost aggregation on the cost volume, obtaining an aggregation result; Using an adaptive multi-modal cross-entropy loss function to supervise the aggregation result, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

2. The method according to claim 1, characterized in that The performing hierarchical feature extraction on the left and right views to obtain a feature map, and compressing the feature map to 32 channels through a feature extraction module incorporating structural information, thereby constructing an initial splicing body includes: Performing downsampling on the input left and right views respectively with a convolutional kernel of a preset size and corresponding strides to reduce the image size; Extract unary features at 1 / 4 resolution using 16 residual layers, namely layer feature maps, layer feature maps and layer feature maps each contain 3 residual layers, and the receptive field is enlarged and semantic information is obtained by increasing the number of channels; Fusing all the feature maps at 1 / 4 resolution to generate a feature map with 320 channels, and constructing an initial splicing body by compressing the feature map to 32 channels.

3. The method according to claim 2, characterized in that, The method extracts unary features at 1 / 4 resolution using 16 residual layers, namely layer feature maps, layer feature maps, and layer feature maps each contain 3 residual layers, and expand the receptive field and obtain semantic information by increasing the number of channels, including: At each convolutional feature extraction layer, calculating the local context structure (LCS) with different dilation rates, increasing the LCS of feature extraction, and fusing multiple local structure feature layers with the original convolutional feature layer. The fused features will learn by themselves to balance the appearance information and the structural information; In the LCS, each channel is used to represent the relationship between a central pixel and its adjacent pixels, and its formula is: Among them, represents the pixel point on the k-th channel of the LCS, is the original convolutional feature, represents the offset of the k-th neighboring point, is a function used to measure the cosine similarity between two vectors. In a convolutional neural network, low-level features contain rich texture information, while high-level features usually reveal semantic clues. By extracting features from different layers, the neighborhood relationship is calculated.

4. The method according to claim 2, wherein The obtaining the attention weights of the corresponding stereo images of the left and right views, using the attention weights to screen the initial splicing body, and constructing a cost volume according to the screening result and the initial disparity loss includes: Step 1: Construction of the initial cost volume; The size of the stereo image pair input to the network is H×W×3. After each image passes through the feature extraction step, unary feature maps of the left view and the right view are obtained respectively. and ; Feature map ( ) has a size of ×H / 4×W / 4, where = 32, representing the number of channels, H: height, W: width, and the initial cost volume is formed by concatenating the feature map and at each disparity level; Step 2: Generation of the attention weights; Obtaining three feature maps at different levels from the feature extraction module, with the number of channels being 64, 128, and 128 respectively; For each pixel at a specific level, using an expansion block with a predefined size and adaptive learning weights to calculate the matching cost; By controlling the dilation rate, ensuring that the range of the window is related to the feature map level, and at the same time, to keep the same number of pixels in calculating the similarity of the central pixel, the similarity of two corresponding pixels is the weighted sum of the correlations of the corresponding pixels within the window; Step 3: Divide the features into groups and calculate the correlation maps for each group; Divide the channels evenly into groups, where the first 8 groups come from layer feature maps, the middle 16 groups come from layer feature maps, and the last 16 groups come from layer feature maps; Feature maps of different levels do not interfere with each other. Denote the g-th group of features as 、 , the matching cost volume of the g-th group, and the matching weight of the g-th group of features, which are adaptively learned during the training process. The final multi-level block matching cost volume is obtained by concatenating the matching costs of all levels; Applying two 3D convolutions and a 3D hourglass network to normalize the obtained multi-level block matching cost volume, and then using another convolutional layer to compress the channels to 1 to obtain the attention weights.

5. The method according to claim 2, characterized in that, The fusing the convolutional context information and the preliminarily aggregated intermediate features to perform cost aggregation on the cost volume, obtaining an aggregation result includes: Given the context features obtained in the convolutional feature extraction step from the left view , and the geometric features obtained by the initial aggregation of the attention cost volume ; Expand the size of to , denoted as , where N represents the number of channels of the feature maps at different levels, and D represents the maximum disparity; Generating spatial attention weights by exploring the spatial relationship between the convolutional context and the geometric features, and using the attention mechanism to fuse the features, thereby adaptively selecting the important regions of the features; Using an aggregation strategy based on the hourglass structure to process the attention cost volume, and realizing multi-scale fusion of the features through progressive downsampling and upsampling.

6. The method according to claim 2, wherein Supervising the aggregation result using the adaptive multi-modal cross-entropy loss function, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views, including: The training process of the network includes: Phase 1: Only train the attention weight part and supervise it using the modeling based on the ground truth probability; Phase 2: Freeze the attention weight part to focus on training other parts of the network; Phase 3: Train the entire network together to achieve the improvement of the overall performance; In each phase, the adaptive multi-modal cross-entropy loss function is used to supervise the aggregation result, and the cross-entropy of this adaptive multi-modal is expressed as: Among them, is the predicted disparity probability distribution obtained through the attention weight A, represents the coefficient of is the i-th predicted disparity probability distribution output during the training process, represents the coefficient of represents the above-mentioned ground truth probability modeling.

7. The method according to claim 6, wherein Supervising the aggregation result using the adaptive multi-modal cross-entropy loss function, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views also includes: Separate each modality in the multi-modal output and calculate the cumulative probability respectively; Each modality represents a potential matching object with a specific depth, and its cumulative probability represents the possibility of the object matching; Adopt the object-level WTA strategy to select the modality with the highest cumulative probability as the dominant modality and normalize this modality: Finally, estimate the disparity using weighted average operation, and estimate the multi-modal output of the network through a multi-modal disparity estimator to obtain the final disparity map.

8. A binocular disparity map acquisition device for small targets, characterized in that Including: An extraction module for performing hierarchical feature extraction on the left and right views to obtain feature maps, and compressing the feature maps to 32 channels through a feature extraction module incorporating structural information, thereby constructing an initial mosaic; A construction module for obtaining the attention weights of the corresponding stereo images of the left and right views, using the attention weights to screen the initial mosaic, and constructing a cost volume according to the screening result and the initial disparity loss; An aggregation module for fusing the convolutional context information and the preliminarily aggregated intermediate features to perform cost aggregation on the cost volume to obtain an aggregation result; A supervision module for supervising the aggregation result using the adaptive multi-modal cross-entropy loss function, and performing multi-modal output on the supervised aggregation result through a multi-modal disparity estimator to obtain the binocular disparity maps of the left and right views.

9. The device according to claim 8, wherein The extraction module includes: A sampling module for downsampling the input left and right views respectively with a corresponding stride through a convolutional kernel of a preset size to reduce the image size; An acquisition module, which is used to extract unary features at a resolution of 1 / 4 by using 16 residual layers, that is, layer feature maps, layer feature maps and layer feature maps each contain 3 residual layers, and the receptive field is expanded and semantic information is obtained by increasing the number of channels; A fusion module for incorporating local structural information LCS into each layer of the convolutional features, fusing all the feature maps at 1 / 4 resolution to generate a feature map with 320 channels, and compressing the feature map to 32 channels, thereby constructing an initial mosaic.

10. A binocular disparity map acquisition system for small targets, characterized in that, Including the device according to any one of claims 8-9.

Citation Information

Patent Citations

  • Target detection method based on improved YOLOv5 and binocular stereo vision

    CN114565900A

  • Multi-modal image registration method based on parallax estimation

    CN115471397A

  • Self-attention multi-scale pyramid binocular stereo matching method and electronic equipment

    CN115861667A

  • Deep stereo matching algorithm based on center pixel gradient fusion and global cost aggregation

    CN115984349A

  • Binocular depth estimation method and system based on attention mechanism and multilevel cost body

    CN116258758A

Cited By

  • Stereo matching method fusing edge guidance and hierarchical cost refinement

    CN121213963A