Stereo matching method and related equipment

By normalizing the pixel eigenvalues of stereo matching images and weighting processing, the accuracy and accuracy problems of stereo matching technology in complex environments are solved, and higher pixel matching capabilities and stability are achieved.

CN120472187APending Publication Date: 2025-08-12BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380819.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing stereo matching technology has reduced pixel matching capabilities under the influence of light, camera imaging quality and scene domain changes, resulting in a decrease in the accuracy and accuracy of stereo matching.

Method used

By normalizing the pixel eigenvalues of multiple images to be matched, especially zero mean normalizing, the pixel matching similarity is determined, and the target cost volume is weighted by using attention weight data to process, cost aggregation and parallax regression are performed to generate body matching results.

Benefits of technology

It improves the accuracy and accuracy of stereo matching, enhances the pixel matching ability, improves the performance and stability of the algorithm, and adapts to the generalization ability of different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472187A_ABST
    Figure CN120472187A_ABST
Patent Text Reader

Abstract

The invention relates to a stereo matching method and related equipment, a vehicle comprises a stereo matching device, the stereo matching device and a computer readable storage medium, and a computer program product is used for realizing the stereo matching method so as to perform normalization processing on pixel characteristic values in a plurality of to-be-matched images. Therefore, the pixel matching similarity among the plurality of to-be-matched images is determined, a stereo matching result is obtained, the distribution difference of pixel features in the plurality of to-be-matched images is reduced, the pixel matching capability is enhanced, and the stereo matching accuracy and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a stereo matching method and related equipment. Background Art

[0002] Stereo matching technology establishes a matching relationship between pixels in the left and right views, thereby determining the disparity between the two views. This disparity can be used to infer depth information in the 3D scene. Related technologies can learn matching rules from the left and right views, achieving end-to-end disparity regression estimation to improve the accuracy and speed of stereo matching.

[0003] However, due to factors such as lighting, camera imaging quality, texture structure, and scene domain changes, the pixel matching capability will decrease, resulting in a decrease in the accuracy and precision of stereo matching. Summary of the Invention

[0004] The embodiments of the present application provide a stereo matching method and related devices to improve the accuracy and precision of stereo matching, so as to at least partially solve the above-mentioned technical problems.

[0005] In order to achieve the above-mentioned object, according to a first aspect of the present application, a stereo matching method is provided, comprising:

[0006] Normalizing the pixel feature values in the multiple images to be matched respectively to obtain multiple processed pixel feature values;

[0007] Based on the multiple processed pixel feature values, pixel matching similarities between the multiple images to be matched are determined, and a stereo matching result is determined according to the pixel matching similarities.

[0008] Optionally, the normalization process includes:

[0009] For each coordinate point in the image to be matched, the processed pixel feature value of the coordinate point is determined according to the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched.

[0010] Optionally, the mean value of the pixel feature of the coordinate point in the image to be matched is determined according to the average of the pixel feature values of multiple feature channels of the coordinate point in the image to be matched.

[0011] Optionally, determining the processed pixel feature value of the coordinate point according to the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched includes:

[0012] For each of the feature channels in the image to be matched, the processed pixel feature value of the coordinate point in the feature channel is determined based on the difference between the pixel feature value of the coordinate point in the feature channel in the image to be matched and the mean value of the pixel features of the coordinate point in the image to be matched.

[0013] Optionally, determining pixel matching similarities between a plurality of images to be matched based on a plurality of processed pixel feature values includes:

[0014] Determining a preset disparity value within a preset disparity range;

[0015] Determining a plurality of associated coordinate points having the preset disparity values in a plurality of images to be matched;

[0016] The pixel matching similarity between the processed pixel feature values of the same feature channel of multiple associated coordinate points is determined.

[0017] Optionally, the pixel matching similarity includes cosine similarity.

[0018] Optionally, determining a stereo matching result according to the pixel matching similarity includes:

[0019] Generate a target cost volume according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities at the associated coordinate points;

[0020] The stereo matching result is determined using the target cost volume.

[0021] Optionally, generating a target cost volume according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities at associated coordinate points includes:

[0022] Determine an initial cost volume under each feature channel according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities under the associated coordinate points;

[0023] The initial cost volume under the feature channel is weighted using the attention weight data under each feature channel to obtain a weighted cost volume under the feature channel, and the target cost volume includes the weighted cost volume under each feature channel.

[0024] Optionally, the attention weight data under each feature channel is determined by the following steps:

[0025] The pixel feature values of the image to be matched in the feature channel are subjected to three-dimensional convolution processing and activation function processing to obtain attention weight data under the feature channel.

[0026] Optionally, determining the stereo matching result by using the target cost volume includes:

[0027] Performing cost aggregation processing on the target cost volume to obtain an aggregated cost volume;

[0028] The aggregated cost volume is upsampled and subjected to disparity regression processing to obtain first disparity maps of multiple images to be matched, and the stereo matching result includes the first disparity maps of multiple images to be matched.

[0029] Optionally, the cost aggregation process includes:

[0030] The target cost volume is subjected to multi-layer downsampling of different scales, multi-layer upsampling of different scales, and residual connection to obtain the aggregated cost volume.

[0031] Optionally, the result obtained after at least one layer of downsampling and / or at least one layer of upsampling of the target cost volume includes pixel matching similarity after weighted processing using attention weight data under the corresponding feature channel.

[0032] Optionally, the disparity regression processing is performed using a disparity regression formula, the disparity regression formula includes a correlation relationship between a predicted disparity value, a preset disparity value, and a target cost volume, and the first disparity map of the multiple images to be matched includes the predicted disparity value.

[0033] Optionally, after determining the stereo matching result using the target cost volume, the method further includes:

[0034] The stereo matching result is used to determine the distance information and / or elevation information in the three-dimensional scene where the multi-camera is located, and multiple images to be matched are obtained based on the multi-camera.

[0035] Optionally, the stereo matching result includes a first disparity map of a plurality of images to be matched, and determining the distance information and / or elevation information in the three-dimensional scene where the multi-camera is located by using the stereo matching result includes:

[0036] Determining three-dimensional information in the three-dimensional scene using first disparity maps of the plurality of images to be matched;

[0037] The distance information and / or elevation information is determined based on the three-dimensional information.

[0038] Optionally, the three-dimensional information is determined using a conversion formula between disparity and three-dimensional coordinates, the conversion formula including the predicted disparity values in the first disparity map of multiple images to be matched, the three-dimensional coordinates in the three-dimensional information, the baseline length of the multi-camera, the camera focal length, and the relationship between the camera principal point coordinates.

[0039] Optionally, after determining the distance information and / or elevation information of the three-dimensional scene where the multi-camera is located using the stereo matching result, the method further includes:

[0040] The distance information and / or elevation information is used to adjust the suspension posture of the vehicle where the multi-camera is located.

[0041] Optionally, the image to be matched is a depth feature map of an initial image to be matched.

[0042] Optionally, the scale of the depth feature map is smaller than the scale of the initial image.

[0043] Optionally, the scale of the depth feature map is one-quarter scale.

[0044] Optionally, the image to be matched is obtained by performing multi-layer downsampling of different scales, multi-layer upsampling of different scales, and residual connection on the initial image.

[0045] Optionally, at least two layers in the multi-layer downsampling and the multi-layer upsampling share weights.

[0046] Optionally, the initial image is an image after epipolar correction processing.

[0047] Optionally, the initial image is determined based on images captured by multiple cameras.

[0048] Optionally, the method is implemented based on a stereo matching network, and the stereo matching network is trained by the following steps:

[0049] Performing zero-mean normalization processing on pixel feature values in a plurality of preset training images to obtain stereo matching results of the plurality of preset training images;

[0050] Determine the target loss function using stereo matching results of multiple preset training images;

[0051] The network model parameters are trained based on the target loss function to obtain the stereo matching network.

[0052] Optionally, the stereo matching results of the plurality of preset training images include first disparity maps of the plurality of preset training images, and determining the target loss function using the stereo matching results of the plurality of preset training images includes:

[0053] The target loss function is determined using the first disparity maps of a plurality of preset training images.

[0054] Optionally, determining the target loss function by using the first disparity maps of a plurality of preset training images includes:

[0055] The target loss function is determined using at least one of the second disparity map and the third disparity map of the plurality of preset training images and the first disparity map of the plurality of preset training images.

[0056] Optionally, the second disparity map is obtained by performing cost aggregation processing and disparity regression processing on a target cost volume in a stereo matching network.

[0057] Optionally, the third disparity map is obtained by upsampling an initial cost volume in a stereo matching network and performing disparity regression processing.

[0058] Optionally, determining the target loss function by using at least one of the second disparity map and the third disparity map of the plurality of preset training images and the first disparity map of the plurality of preset training images includes:

[0059] The target loss function is obtained by performing weighted summation using the loss function of at least one of the second disparity map and the third disparity map and the loss functions of the first disparity maps of multiple preset training images.

[0060] Optionally, the weight of the loss function of the first disparity map of the plurality of preset training images is greater than the weight of the loss function of at least one of the second disparity map and the third disparity map.

[0061] According to a second aspect of the present application, a stereo matching device is provided, comprising a processor, a memory, and a computer program or instructions, wherein the computer program or instructions are stored in the memory and executed by the processor to implement any one of the stereo matching methods described.

[0062] According to a fifth aspect of the present application, a computer-readable storage medium is provided, on which a computer program or instructions are stored, and the computer program or instructions are executed by a processor to implement any one of the stereo matching methods described.

[0063] According to a sixth aspect of the present application, a computer program product is provided, comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement any one of the stereo matching methods described.

[0064] According to a seventh aspect of the present application, a vehicle is provided, comprising any one of the stereo matching devices described.

[0065] In summary, the vehicle of the embodiment of the present application includes a stereo matching device, and the stereo matching device, a computer-readable storage medium, and a computer program product are used to implement a stereo matching method, thereby normalizing the pixel feature values in multiple images to be matched, and then determining the pixel matching similarity between the multiple images to be matched, obtaining a stereo matching result, reducing the distribution differences of pixel features in the multiple images to be matched, enhancing the pixel matching capability, and improving the accuracy and precision of stereo matching.

[0066] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0068] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.

[0069] Figure 1 is a flowchart of the steps of the stereo matching method provided in an exemplary embodiment of the present application;

[0070] Figure 2 is a schematic structural diagram of a stereo matching network provided in an exemplary embodiment of the present application;

[0071] Figure 3 is a schematic diagram of weighted processing of attention weight data provided in an exemplary embodiment of the present application;

[0072] Figure 4 is a schematic diagram of an initial image provided in an exemplary embodiment of the present application;

[0073] Figure 5 is based on Figure 4 A schematic diagram of a first disparity map determined from an initial image;

[0074] Figure 6 is based on Figure 5 A schematic diagram of a depth map of distance information determined by a first disparity map;

[0075] Figure 7 is based on Figure 6 A schematic diagram of the RGBD three-dimensional point cloud determined by the depth map;

[0076] Figure 8This is an application scenario diagram of the vehicle-mounted multi-eye preview system provided in an exemplary embodiment of the present application;

[0077] Figure 9 It is a structural diagram of a stereo matching device provided in an exemplary embodiment of the present application.

[0078] in, Figure 2 、 4 , 5, 6, and 7 are color images, which reflect the true status of the image at different processing stages through different colors. DETAILED DESCRIPTION

[0079] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0080] Based on the issues mentioned in the previous background, stereo matching technology, a related technique, establishes a matching relationship between pixels in a pair of left and right views after epipolar correction, and then uses triangulation to estimate the depth information of the three-dimensional scene. Stereo matching is a key technology in multi-view (e.g., at least two-view) depth estimation tasks. It can be divided into stereo matching methods based on manual geometric features and end-to-end stereo matching methods based on deep learning.

[0081] Stereo matching methods based on manual geometric features rely on pre-set feature construction methods to obtain the similarity relationship of pixel points. This method can obtain more accurate depth information, but it is still affected by occlusion, featureless areas or areas with repeated textures, which limits the performance of this method in complex scenes.

[0082] End-to-end stereo matching methods based on deep learning, by training network model parameters, enable the neural network to learn pixel matching rules from the left and right views, achieving end-to-end disparity regression estimation, significantly improving accuracy and speed. This end-to-end stereo matching method based on deep learning can be roughly divided into four steps: deep feature extraction, target cost volume (CV) construction, cost aggregation, and disparity regression. The target cost volume, constructed by introducing geometric prior knowledge within a preset disparity range, is a module used in stereo matching to measure the similarity of pixel matching between the left and right views and plays a crucial role in multi-view depth estimation. The target cost volume construction method uses operations such as feature concatenation, feature cosine similarity, and feature difference in the disparity direction to establish the disparity relationship between the left and right views. This method often achieves better accuracy. However, due to factors such as ambient lighting, scene structure, and camera image quality, the matching performance of the target cost volume constructed from the depth feature maps extracted from the left and right views can be reduced. The stereo matching network can improve the quality of the target cost volume by weighting the target cost volume or improving the regularization ability during cost aggregation. However, these methods will increase the network model parameters of the stereo matching network, and a more complex stereo matching network will reduce the efficiency of stereo matching.

[0083] To this end, the embodiment of the present application analyzes the principle of pixel matching of the left and right views achieved by the target cost volume and the limitations of the cost volume construction method, and designs a target cost volume constructed based on pixel matching similarity of Z-Score (zero mean normalization), as well as a stereo matching network that uses intermediate supervision to improve the quality of the target cost volume.

[0084] According to a first aspect of the present application, a stereo matching method is provided, corresponding to Figure 1 According to the content of "Constructing Correlation Cost Volume Based on Z-Score", the stereo matching method can include the following steps:

[0085] Normalization is performed on pixel feature values in multiple images to be matched to obtain multiple processed pixel feature values; based on the multiple processed pixel feature values, pixel matching similarities between the multiple images to be matched are determined, and a stereo matching result is determined according to the pixel matching similarities.

[0086] In the embodiment of the present application, the multiple images to be matched are multiple images that need to be stereo matched. For example, the multiple images to be matched can be the left and right views (i.e. Figure 1 The "left and right images" involved in the left view or right view can be specifically, for example Figure 4), or images obtained by further processing the left and right views (for example, images obtained by preprocessing and epipolar correction of the left and right views, followed by deep feature extraction, where the preprocessing may be, for example, filtering). The multi-camera includes at least two cameras, for example, a binocular camera, a trinocular camera, etc.; accordingly, the multiple images to be matched can be acquired based on at least two of the multi-cameras, for example, a binocular camera, or a combination of data acquired based on at least two of the trinocular cameras. The pixel feature value can be a pixel grayscale value, a pixel RGB (red, green, blue) value, etc.

[0087] Normalization is the process of limiting pixel feature values to a certain range after data processing. For example, Z-Score (zero-mean normalization) can be used. Zero-mean normalization reduces the difference between pixel feature values by subtracting their corresponding pixel feature mean, thereby enhancing pixel matching capabilities during stereo matching.

[0088] As can be seen, through the above technical solution, the pixel feature values in multiple images to be matched are respectively normalized, and then the pixel matching similarity between the multiple images to be matched is determined to obtain a stereo matching result, thereby reducing the distribution differences of the pixel features in the multiple images to be matched, enhancing the pixel matching capability, and improving the accuracy and precision of stereo matching. In addition, it should be noted that the pixel matching similarity can describe the directional correlation of the pixel features in the multiple images to be matched. The embodiment of the present application can eliminate the different dimensions of the pixel features in the multiple images to be matched and convert them to the same standard scale for pixel matching, thereby improving the performance and stability of the algorithm and making the pixel matching more domain-invariant.

[0089] In some embodiments, taking zero-mean normalization as an example, the normalization process may specifically include: for each coordinate point in the image to be matched, determining a processed pixel feature value for the coordinate point based on the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched. For example, the difference between the pixel feature value of each coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched may be directly used as the processed pixel feature value of the coordinate point, thereby achieving zero-mean normalization for the image to be matched.

[0090] In some embodiments, the image to be matched may include multiple feature channels (channel, C), and the feature channels may be, for example, the R (red) feature channel, the G (green) feature channel, the B (blue) feature channel, etc. of the image to be matched. Since the pixel feature values of each coordinate point in the image to be matched in different feature channels may be different, a pixel feature mean may be determined for each coordinate point in the image to be matched. Specifically, the pixel feature mean of each coordinate point in the image to be matched may be determined based on the average of the pixel feature values of the coordinate point in multiple feature channels in the image to be matched. For example, the average of the pixel feature values of the coordinate point in multiple feature channels in the image to be matched may be directly used as the pixel feature mean of the coordinate point in the image to be matched, thereby determining a more representative pixel feature mean. The formula for the pixel feature mean of the coordinate point (x, y) in the image to be matched may be, for example:

[0091] μ L (:, x, y) = ∑ K∈[0,C] f L (K, x, y) / C

[0092] μ R (:, x, y) = ∑ K∈[0,C] f R (K, x, y) / C

[0093] Among them, μ L (:, x, y) is the pixel feature mean of the coordinate point (x, y) of the left view in multiple images to be matched, C is the feature channel, f L (K, x, y) is the pixel feature value of the coordinate point (x, y) in the left view in the Kth feature channel. R (:, x, y) is the pixel feature mean of the coordinate point (x, y) of the right view in multiple images to be matched, C is the feature channel, f R (K, x, y) is the pixel feature value of the coordinate point (x, y) of the right view in the Kth feature channel.

[0094] In some embodiments, since the image to be matched may include multiple feature channels, determining the processed pixel feature value of the coordinate point based on the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched may include: for each feature channel in the image to be matched, determining the processed pixel feature value of the coordinate point in the feature channel based on the difference between the pixel feature value of the coordinate point in the feature channel in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched. For example, the difference between the pixel feature value of the coordinate point in the feature channel in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched may be directly used as the processed pixel feature value of the coordinate point in the feature channel, thereby distinguishing different feature channels and achieving zero-mean normalization processing for each pixel in each feature channel of the image to be matched, eliminating scale and position differences of pixel features, and at the same time enabling the stereo matching method to have better generalization for unlearned scene domains, thereby improving the three-dimensional perception capability of the scene.

[0095] In some embodiments, an example is provided for determining pixel matching similarity. Specifically, based on multiple processed pixel feature values, determining the pixel matching similarity between multiple images to be matched may include: determining a preset disparity value within a preset disparity range, for example, determining a preset disparity value d within a preset disparity range [0, D); determining multiple associated coordinate points with preset disparity values in the multiple images to be matched, for example, where the multiple processed images include a processed left view and a processed right view, a first coordinate point (x, y) may be determined in the processed left view, and then a second coordinate point (xd, y) having a preset disparity value d with the first coordinate point may be determined in the processed right view, where the multiple associated coordinate points include the first coordinate point and the second coordinate point; and determining the pixel matching similarity between the processed pixel feature values of the multiple associated coordinate points in the same feature channel, thereby calculating the pixel matching similarity.

[0096] In some embodiments, the pixel matching similarity may specifically include cosine similarity. Cosine similarity evaluates the similarity of two vectors by calculating the cosine value of the angle between them. Therefore, in the embodiment of the present application, it can be used to evaluate the similarity of the processed pixel feature values of multiple associated coordinate points in the same feature channel to enhance the pixel matching ability during stereo matching. Of course, in other embodiments, pixel matching similarity may also include Jaccard similarity, Pearson similarity, etc. The formula for cosine similarity can be, for example:

[0097]

[0098] Among them, V(:,D / 4,x,y) refers to the cosine similarity between the processed pixel feature values of multiple associated coordinate points in the same feature channel, fL (:, x, y) is the pixel feature value of the coordinate point (x, y) of the left view in the multiple images to be matched in the feature channel, f r (:, xd, y) represents the pixel feature value of the coordinate point (x, y) in the right view of the multiple images to be matched in this feature channel. D represents the boundary value of the preset disparity range [0, D). The meaning of D / 4 is described below.

[0099] In some embodiments, determining a stereo matching result based on pixel matching similarity may include: generating a target cost volume based on pixel matching similarities of multiple images to be matched, and / or preset disparity values (d), and / or feature channels (C), and / or associated coordinate points (x, y); and determining the stereo matching result using the target cost volume. It can be seen that the pixel matching similarities of multiple images to be matched, preset disparity values (d), feature channels (C), associated coordinate points (x, y), and other dimensions are used to construct the target cost volume to enhance the pixel matching capability of the target cost volume. The multiple images to be matched used to generate the target cost volume may be images from a single batch (batch, B) (e.g., left and right views), i.e., the target cost volume is generated based on multiple images to be matched from the same batch. Of course, the multiple images to be matched used to generate the target cost volume may also be images from multiple batches (e.g., left and right views), i.e., the target cost volume is generated based on multiple images to be matched from different batches, without limitation herein.

[0100] In some embodiments, reference Figure 2 and Figure 3 In order to enhance the geometric expression ability of the target cost volume, the target cost volume can also be constructed by weighted processing of the attention weight data. Specifically, according to the pixel matching similarity of multiple images to be matched, and / or preset disparity values, and / or feature channels, and / or associated coordinate points, generating the target cost volume can include: determining the initial cost volume C under each feature channel according to the pixel matching similarity of multiple images to be matched, and / or preset disparity values, and / or feature channels, and / or associated coordinate points. i , the initial cost volume under each feature channel may include multiple images to be matched, and / or preset disparity values, and / or pixel matching similarities under associated coordinate points. Therefore, the size of the initial cost volume under each feature channel can be, for example, B*D / 4*H*W. The initial cost volume is a 4-dimensional cost volume, B is the batch, C is the feature channel, H is the image height, and W is the image width. Using the attention weight data (ATT, Attetion) under each feature channel, the initial cost volume C under the feature channel is iPerform weighted processing to obtain the weighted cost volume C0 under the feature channel. The target cost volume may include the weighted cost volume under each feature channel, that is, the size of the target cost volume may be, for example, B*C*D / 4*H*W, and the target cost volume is a 5-dimensional cost volume, thereby achieving dimensionality increase of the cost volume (i.e., from the 4-dimensional initial cost volume to the 5-dimensional target cost volume, to enrich the feature extraction and pixel matching capabilities of the target cost volume).

[0101] It can also be seen that weighted processing of attention weight data can stimulate the relevant geometric features of the corresponding initial cost volume, thereby enhancing the geometric expression ability of the target cost volume. Without performing neighborhood aggregation, weighted processing of attention weight data can make the spatial variation of the target cost volume more efficient.

[0102] In some embodiments, the attention weight data under each feature channel can be determined by the following steps: the pixel feature values of the matching image in the feature channel are subjected to three-dimensional convolution (CNN, Convolutional Neural Networks) processing and activation function (e.g. Figure 3 The sigmoid function in [ ] is used to process the attention weight data (ATT, Attition) for that feature channel. As can be seen, the attention weight data for each feature channel records the pixel feature information of the image to be matched in that feature channel, thus accurately stimulating the relevant geometric features of the corresponding initial cost volume, improving the accuracy and precision of stereo matching.

[0103] In some embodiments, determining the stereo matching result using the target cost volume may include: performing cost aggregation processing (i.e., regularization processing) on the target cost volume to obtain an aggregated cost volume, so that the aggregated cost volume can obtain more global context information. The size of the aggregated cost volume may be, for example, B*D / 4*H*W, and the target cost volume is a 4-dimensional cost volume; performing upsampling and disparity regression processing on the aggregated cost volume to obtain a first disparity map (i.e., Figure 1 In the "Disparity Map No. 2", Figure 2 The stereo matching results include first disparity maps for multiple images to be matched, each of which includes the predicted disparity values for each coordinate point in the images to be matched, thereby achieving stereo matching for the multiple images to be matched. Upsampling ensures that the scale of the final output first disparity map is the same as the scale of the original input image from the multi-camera.

[0104] In some embodiments, the cost aggregation process is illustrated. Specifically, the cost aggregation process includes: performing multi-layer downsampling of different scales, multi-layer upsampling of different scales, and residual connection (skip connection) on the target cost volume to obtain an aggregated cost volume. Among them, the multiple different scales can be, for example, a quarter scale, an eighth scale, a sixteenth scale, a thirty-second scale, etc., and the multi-layer downsampling can be achieved by a multi-layer encoder, and the multi-layer upsampling and residual connection can be achieved by a multi-layer decoder. The multi-layer encoder and the multi-layer decoder can constitute a coding-decoding network. The structural type of the coding-decoding network can be, for example, U-Net, etc., so as to retain the structural information and low-frequency information in the image as much as possible. The coding-decoding network can be, for example, MobileNet, EfficientNet, etc.

[0105] In some embodiments, at least two layers of multi-layer downsampling and multi-layer upsampling share weights. By sharing weights, the number of network model parameters in the encoding-decoding network can be reduced, thereby improving the speed of stereo matching.

[0106] In some embodiments, in order to enhance the geometric expression ability of the target cost volume, the step of cost aggregation processing on the target cost volume can also be combined with the weighted processing of attention weight data. Figure 2 , the result obtained after at least one layer of downsampling and / or at least one layer of upsampling of the target cost volume includes the pixel matching similarity after weighted processing using the attention weight data under the corresponding feature channel. For example, in the process of performing multi-layer downsampling of different scales, multi-layer upsampling of different scales and residual connection on the target cost volume, the result obtained after each layer of downsampling, each upsampling and residual connection can be weighted using the attention weight data of the corresponding scale under the corresponding feature channel, so that the pixel matching similarity in the result obtained after each layer of downsampling, each upsampling and residual connection is weighted by the attention weight data of the corresponding scale under the corresponding feature channel. In this way, the obtained aggregated cost volume can include more texture information in the image to be matched, thereby enhancing the geometric expression ability of the aggregated cost volume.

[0107] In some embodiments, the disparity regression process can be implemented using a disparity regression formula. The disparity regression formula includes the relationship between the predicted disparity value, the preset disparity value, and the target cost volume. For example, the disparity regression formula can be:

[0108]

[0109] in, is the predicted disparity value, d is the preset disparity value, and d is a value in the preset disparity range [0, D), softmax is the index of the normalized element, c d is the aggregate cost volume.

[0110] In some embodiments, after determining the stereo matching result using the target cost volume, the method further includes: determining the distance information and / or elevation information (the depth map of the distance information, for example) in the three-dimensional scene where the multi-camera is located using the stereo matching result. Figure 6 As shown, the RGBD three-dimensional point cloud of the three-dimensional scene is as follows: Figure 7 As shown). It can be seen that by determining the distance information and / or elevation information in the three-dimensional scene, the stereo matching results can be better applied to the actual scene.

[0111] In some embodiments, the stereo matching result includes a first disparity map of multiple images to be matched (eg Figure 5 Taking the example of FIG. 1 as shown in FIG. 2 , using the stereo matching results to determine the distance information and / or elevation information in the three-dimensional scene where the multi-camera is located can include: using the first disparity maps of the multiple images to be matched to determine the three-dimensional information in the three-dimensional scene, where the three-dimensional information may include, for example, the three-dimensional coordinates of each position in the three-dimensional scene where the multi-camera is located; and determining the distance information and / or elevation information based on the three-dimensional information, so that the determination of the distance information and / or elevation information is more convenient.

[0112] In some embodiments, an example method for determining three-dimensional information is provided. Specifically, the three-dimensional information can be determined using a conversion formula between disparity and three-dimensional coordinates. The conversion formula includes the relationship between the predicted disparity values in the first disparity map of the multiple images to be matched, the three-dimensional coordinates in the three-dimensional information, the baseline length of the multi-camera, the camera focal length, and the coordinates of the camera principal point. For example, the conversion formula can be:

[0113]

[0114] Where X, Y, and Z are the 3D coordinates in the 3D information, d is the disparity in the first disparity map, B is the baseline length of the multi-camera, f is the camera focal length, and (x0, y0) are the coordinates of the camera's principal point. This conversion formula allows for rapid conversion of the first disparity maps of multiple images to be matched into the 3D coordinates of each location in the 3D scene where the multi-camera resides.

[0115] In some embodiments, the distance information and / or elevation information of the three-dimensional scene where the multi-camera is located can be used to adjust the vehicle suspension posture. Specifically, after using the stereo matching results to determine the distance information and / or elevation information of the three-dimensional scene where the multi-camera is located, the method may further include: using the distance information and / or elevation information to adjust the suspension posture of the vehicle where the multi-camera is located. For example Figure 8The vehicle-mounted multi-eye preview system (for example, a vehicle-mounted binocular preview system) shown is a system that collects road surface information ahead through the vehicle's image sensor and provides it to the vehicle's suspension controller to adjust the suspension settings in advance, allowing the vehicle to smoothly pass over uneven roads. In the vehicle-mounted multi-eye preview system, the vehicle where the multi-eye camera is located can use the distance information and / or elevation information in the three-dimensional scene where the multi-eye camera is located to perceive irregular roads or obstacles, identify obstructed roads, and then use the suspension controller to adjust the suspension settings in advance (adjust the damping state of the vehicle's suspension shock absorber) so that the vehicle can smoothly pass over obstructed roads.

[0116] In some embodiments, the image to be matched is a depth feature map of the initial image to be matched, that is, Figure 1 and Figure 2 In the steps, a corresponding depth feature map is obtained by extracting the depth feature of the initial image and serving as the corresponding image to be matched, so that the image to be matched includes more image features to improve the accuracy and precision of stereo matching.

[0117] In some embodiments, the scale of the depth feature map is smaller than the scale of the initial image, that is, the scale of the image to be matched is smaller than the scale of the initial image. In this way, the amount of computation when using the image to be matched for stereo matching will be smaller, thereby improving the speed of stereo matching.

[0118] In some embodiments, the scale of the depth feature map is one-quarter scale, that is, the scale of the image to be matched is one-quarter scale (i.e., D / 4). In this way, the amount of computation when using the image to be matched for stereo matching will be smaller, and the accuracy of stereo matching will not decrease too much, thereby balancing the accuracy and speed of stereo matching.

[0119] In some embodiments, the process of obtaining the image to be matched is illustrated. Figure 2 , the image to be matched can be obtained by comparing the initial image ( Figure 2 The "left image" and "right image" in the figure are obtained by performing multi-layer downsampling at different scales, multi-layer upsampling at different scales, and residual connections (skip connection). Among them, the multi-layer downsampling can be achieved by a multi-layer encoder, and the multi-layer upsampling and residual connections can be achieved by a multi-layer decoder. The multi-layer encoder and the multi-layer decoder can form an encoding-decoding network and are used to extract deep features of the initial image. The structural type of the encoding-decoding network can be, for example, U-Net, etc., and the encoding-decoding network can be, for example, MobileNet, EfficientNet, etc.

[0120] In some embodiments, at least two layers of multi-layer downsampling and multi-layer upsampling share weights. By sharing weights, the number of network model parameters in the encoding-decoding network can be reduced, thereby improving the speed of stereo matching.

[0121] In some embodiments, reference Figure 1 The initial image is an image that has undergone epipolar correction. This means that epipolar correction is performed before deep feature extraction. Epipolar correction using pre-calibrated parameters can make the epipolar lines of different initial images parallel and horizontally aligned. This allows for faster stereo matching when constructing the target cost volume using the image to be matched.

[0122] In some embodiments, the initial image is determined based on images captured by multiple cameras. For example, the images captured by the multiple cameras may be preprocessed and epipolar corrected to obtain the initial image. Preprocessing may include filtering, for example. As can be seen, by determining the initial image based on images captured by multiple cameras, the initial image required for stereo matching can be obtained.

[0123] In some embodiments, the stereo matching method is implemented based on a stereo matching network. The structure of the stereo matching network is as follows: Figure 2 As shown in , the stereo matching network may include a target cost volume. The stereo matching network can be trained by the following steps: performing zero-mean normalization processing on the pixel feature values in multiple preset training images to obtain stereo matching results of the multiple preset training images, wherein the preset training images are images in the training data set used to train the stereo matching network. The method for determining the stereo matching results of the multiple preset training images can refer to the method for determining the stereo matching results of the multiple images to be matched, which is not described in detail here; using the stereo matching results of the multiple preset training images, determining the target loss function; training the network model parameters based on the target loss function to obtain the stereo matching network, for example, by recursively deriving the target loss function to train the network model parameters, and finally obtaining the stereo matching network.

[0124] In some embodiments, the stereo matching results of the multiple preset training images include first disparity maps of the multiple preset training images. The method for determining the first disparity maps of the multiple preset training images can refer to the method for determining the first disparity maps of the multiple images to be matched, and is not further described here. Determining the target loss function using the stereo matching results of the multiple preset training images may include determining the target loss function using the first disparity maps of the multiple preset training images.

[0125] In some embodiments, when determining the target loss function, multiple different disparity maps can also be combined. Specifically, using the first disparity maps of multiple preset training images to determine the target loss function can include: using at least one of the second disparity maps and third disparity maps of multiple preset training images, and the first disparity map of multiple preset training images to determine the target loss function. It can be seen that at least one of the second disparity map and the third disparity map of the multiple preset training images is used to perform intermediate supervision on the training of the target cost volume in the stereo matching network to improve the quality of the target cost volume, so that the stereo matching accuracy and precision of the trained stereo matching network are also higher.

[0126] In some embodiments, the second disparity map can be obtained by performing cost aggregation processing and disparity regression processing on the target cost volume in the stereo matching network (the second disparity map is Figure 1 In the "Disparity Map No. 1", Figure 2 ). As can be seen, the second disparity map is determined differently from the first, and therefore also differs from the first, allowing the target loss function to consider more dimensions. Furthermore, since the second disparity map is not upsampled, its scale is the same as that of the image to be matched; for example, it can be a quarter of the scale.

[0127] In some embodiments, the third disparity map can be obtained by upsampling the initial cost volume in the stereo matching network and performing disparity regression processing (the third disparity map is Figure 1 In the "Disparity Map No. 0", Figure 2 ). As can be seen, the third disparity map is determined differently from the first and second disparity maps, and therefore also differs from them, allowing the target loss function to be considered from multiple dimensions. Furthermore, since the third disparity map is not upsampled, the scale of the second disparity map can be the same as the original input image from the multi-camera.

[0128] In some embodiments, determining a target loss function using at least one of the second disparity map and the third disparity map of multiple preset training images, and the first disparity map of multiple preset training images, may include: performing weighted summation using the loss function of at least one of the second disparity map and the third disparity map, and the loss function of the first disparity map of multiple preset training images to obtain the target loss function, thereby avoiding the lack of supervisory information constraints during the stereo matching network training process, thereby causing the entire network to overfit, and improving the pixel matching capability of the stereo matching network. The formula of the target loss function may be, for example:

[0129] Loss total=α·Loss0+β·Loss1+γ·Loss2

[0130] Among them, Loss total is the target loss function, Loss0 is the loss function for the third disparity map, Loss1 is the loss function for the second disparity map, and Loss2 is the loss function for the first disparity map of multiple preset training images. α, β, and γ are the weights of Loss0, Loss1, and Loss2, respectively. The value ranges of α, β, and γ can be, for example, 0 < α ≤ 0.5, 0 < β ≤ 0.5, 1 ≤ γ ≤ 1.5, and α + β + γ = 1.5.

[0131] In some embodiments, by adjusting the weight of the loss function, the degree of emphasis that the stereo matching network places on a certain disparity map can be effectively enhanced. For example, since the first disparity map of multiple preset training images is the final output result of the stereo matching network, and the second disparity map and the third disparity map are used for training the stereo matching network, when the stereo matching network is applied (for example, using the stereo matching network for model inference), the second disparity map and the third disparity map are not generated. Therefore, the weight of the loss function of the first disparity map of the multiple preset training images may be greater than the weight of the loss function of at least one of the second disparity map and the third disparity map, that is, γ>α, and γ>β, for example, α=0.2, β=0.3, and γ=1.

[0132] In addition, since the second disparity map and the third disparity map are not generated when the stereo matching network is used for model inference, the inference burden of the stereo matching network will not be increased, and the edge device does not need to invest more computing resources.

[0133] In some embodiments, the loss function of at least one of the second disparity map and the third disparity map, and at least one of the loss functions of the first disparity map of the plurality of preset training images may be calculated using a smoothed mean absolute error. The formula for the smoothed mean absolute error may be, for example:

[0134]

[0135] in,

[0136] Among them, δ can be 1, N is the total number of pixels with valid labels, d is the true disparity value, is the predicted disparity value.

[0137] According to a second aspect of the present application, a stereo matching device is provided, comprising a processor, a memory, and a computer program or instructions, wherein the computer program or instructions are stored in the memory and executed by the processor to implement any of the stereo matching methods described above. The stereo matching device has all the beneficial effects of the aforementioned stereo matching methods, which are not further described in this application.

[0138] In some embodiments, as Figure 9 , which shows a schematic structural diagram of a stereo matching device according to an embodiment of the present application, specifically:

[0139] The stereo matching device may include one or more processing core processors 901, one or more computer readable storage media memories 902, a power supply 903, an input unit 904 and other components. Those skilled in the art will understand that Figure 9 The structure of the stereo matching device shown in the figure is not intended to limit the stereo matching device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0140] Processor 901 is the control center of the stereo matching device. It uses various interfaces and lines to connect the various parts of the entire stereo matching device. By running or executing software programs and / or modules stored in memory 902 and calling data stored in memory 902, it performs various functions of the stereo matching device and processes data, thereby monitoring the stereo matching device as a whole. Optionally, processor 901 may include one or more processing cores; preferably, processor 901 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly handles wireless communications. It is understood that the above-mentioned modem processor may not be integrated into processor 901.

[0141] The memory 902 can be used to store software programs and modules. The processor 901 executes various functional applications and data processing by running the software programs and modules stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the stereo matching device, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 902 may also include a memory controller to provide the processor 901 with access to the memory 902.

[0142] The stereo matching device also includes a power supply 903 for supplying power to various components. Preferably, the power supply 903 can be logically connected to the processor 901 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 903 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0143] The stereo matching device may further include an input unit 904, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0144] Although not shown, the stereo matching device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 901 in the stereo matching device loads the executable files corresponding to one or more application processes into the memory 902 according to the following instructions. The processor 901 then runs the application stored in the memory 902 to implement various functions, such as implementing any of the steps in the stereo matching method described above.

[0145] According to a third aspect of the present application, a computer-readable storage medium is provided. The storage medium may include a read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk. A computer program or instructions are stored on the storage medium, which is loaded by a processor to execute the steps of any of the above-described stereo matching methods. The computer-readable storage medium has all the advantages of the above-described stereo matching methods, and this application will not further elaborate on them.

[0146] According to a fourth aspect of the present application, a computer program product is provided, comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement any one of the stereo matching methods described above. The computer program product has all the beneficial effects of the stereo matching method described above, and this application will not further elaborate on them.

[0147] According to a fifth aspect of the present application, a vehicle is provided, comprising any of the stereo matching devices described above. The stereo matching device may, for example, be integrated into a domain controller in the vehicle, or include a domain controller in the vehicle. The vehicle has all the beneficial effects of the stereo matching devices described above, which are not further detailed in this application.

[0148] In some embodiments, the vehicle may be a fuel vehicle, a plug-in hybrid vehicle, or a new energy vehicle, etc., and this application does not make any specific limitations on this.

[0149] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0150] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0151] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0152] The above are merely preferred embodiments of the present application and do not constitute any form of limitation to the present application. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A stereo matching method, characterized in that: include: Normalizing the pixel feature values in the multiple images to be matched respectively to obtain multiple processed pixel feature values; Based on the multiple processed pixel feature values, pixel matching similarities between the multiple images to be matched are determined, and a stereo matching result is determined according to the pixel matching similarities.

2. The stereo matching method according to claim 1, wherein: The normalization process includes: For each coordinate point in the image to be matched, the processed pixel feature value of the coordinate point is determined according to the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched.

3. The stereo matching method according to claim 2, wherein: The pixel feature mean of the coordinate point in the image to be matched is determined according to the average of the pixel feature values of the coordinate point in multiple feature channels in the image to be matched.

4. The stereo matching method according to claim 3, wherein: The step of determining the processed pixel feature value of the coordinate point according to the difference between the pixel feature value of the coordinate point in the image to be matched and the mean value of the pixel feature of the coordinate point in the image to be matched comprises: For each of the feature channels in the image to be matched, the processed pixel feature value of the coordinate point in the feature channel is determined based on the difference between the pixel feature value of the coordinate point in the feature channel in the image to be matched and the mean value of the pixel features of the coordinate point in the image to be matched.

5. The stereo matching method according to claim 1, wherein: The determining of pixel matching similarities between the plurality of images to be matched based on the plurality of processed pixel feature values includes: Determining a preset disparity value within a preset disparity range; Determining a plurality of associated coordinate points having the preset disparity values in a plurality of images to be matched; The pixel matching similarity between the processed pixel feature values of the same feature channel of multiple associated coordinate points is determined.

6. The stereo matching method according to claim 5, wherein: The pixel matching similarity includes cosine similarity.

7. The stereo matching method according to claim 5, wherein: Determining a stereo matching result according to the pixel matching similarity includes: Generate a target cost volume according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities at the associated coordinate points; The stereo matching result is determined using the target cost volume.

8. The stereo matching method according to claim 7, wherein: Generating a target cost volume according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities at associated coordinate points, includes: Determine an initial cost volume under each feature channel according to the plurality of images to be matched, and / or the preset disparity values, and / or the feature channels, and / or the pixel matching similarities under the associated coordinate points; The initial cost volume under the feature channel is weighted using the attention weight data under each feature channel to obtain a weighted cost volume under the feature channel, and the target cost volume includes the weighted cost volume under each feature channel.

9. The stereo matching method according to claim 8, wherein: The attention weight data under each feature channel is determined by the following steps: The pixel feature values of the image to be matched in the feature channel are subjected to three-dimensional convolution processing and activation function processing to obtain attention weight data under the feature channel.

10. The stereo matching method according to claim 7, wherein: Determining the stereo matching result by using the target cost volume includes: Performing cost aggregation processing on the target cost volume to obtain an aggregated cost volume; The aggregated cost volume is upsampled and subjected to disparity regression processing to obtain first disparity maps of multiple images to be matched, and the stereo matching result includes the first disparity maps of multiple images to be matched.

11. The stereo matching method according to claim 10, wherein: The cost aggregation process includes: The target cost volume is subjected to multi-layer downsampling of different scales, multi-layer upsampling of different scales, and residual connection to obtain the aggregated cost volume.

12. The stereo matching method according to claim 11, wherein: The result obtained after at least one layer of downsampling and / or at least one layer of upsampling of the target cost volume includes pixel matching similarity after weighted processing using attention weight data under the corresponding feature channel.

13. The stereo matching method according to claim 10, wherein: The disparity regression process is performed using a disparity regression formula, which includes a correlation relationship between a predicted disparity value, a preset disparity value, and a target cost volume, and the first disparity map of the plurality of images to be matched includes the predicted disparity value.

14. The stereo matching method according to claim 7, wherein: After determining the stereo matching result using the target cost volume, the method further includes: The stereo matching result is used to determine the distance information and / or elevation information in the three-dimensional scene where the multi-camera is located, and multiple images to be matched are obtained based on the multi-camera.

15. The stereo matching method according to claim 14, wherein: The stereo matching result includes a first disparity map of a plurality of images to be matched, and the determining of distance information and / or elevation information in a three-dimensional scene where the multi-camera is located by using the stereo matching result includes: Determining three-dimensional information in the three-dimensional scene using first disparity maps of the plurality of images to be matched; The distance information and / or elevation information is determined based on the three-dimensional information.

16. The stereo matching method according to claim 15, wherein: The three-dimensional information is determined using a conversion formula between disparity and three-dimensional coordinates, wherein the conversion formula includes the predicted disparity values in the first disparity map of multiple images to be matched, the three-dimensional coordinates in the three-dimensional information, the baseline length of the multi-camera, the camera focal length, and the relationship between the camera principal point coordinates.

17. The stereo matching method according to claim 14, wherein: After determining the distance information and / or elevation information of the three-dimensional scene where the multi-camera is located using the stereo matching result, the method further includes: The distance information and / or elevation information is used to adjust the suspension posture of the vehicle where the multi-camera is located.

18. The stereo matching method according to claim 1, wherein: The image to be matched is a depth feature map of the initial image to be matched.

19. The stereo matching method according to claim 18, wherein: The scale of the depth feature map is smaller than that of the initial image.

20. The stereo matching method according to claim 19, wherein: The scale of the depth feature map is one-quarter.

21. The stereo matching method according to claim 18, wherein: The image to be matched is obtained by performing multi-layer downsampling of different scales, multi-layer upsampling of different scales and residual connection on the initial image.

22. The stereo matching method according to claim 21, wherein: At least two layers in multi-layer downsampling and multi-layer upsampling share weights.

23. The stereo matching method according to claim 18, wherein: The initial image is the image after epipolar correction.

24. The stereo matching method according to claim 18, wherein: The initial image is determined based on images captured by multiple cameras.

25. The stereo matching method according to any one of claims 1 to 24, wherein: The method is implemented based on a stereo matching network, and the stereo matching network is trained by the following steps: Performing zero-mean normalization processing on pixel feature values in a plurality of preset training images to obtain stereo matching results of the plurality of preset training images; Determine the target loss function using stereo matching results of multiple preset training images; Network model parameter training is performed based on the target loss function to obtain the stereo matching network.

26. The stereo matching method according to claim 25, wherein: The stereo matching results of the plurality of preset training images include first disparity maps of the plurality of preset training images, and determining the target loss function using the stereo matching results of the plurality of preset training images includes: The target loss function is determined using the first disparity maps of a plurality of preset training images.

27. The stereo matching method according to claim 26, wherein: The determining the target loss function by using the first disparity maps of the plurality of preset training images includes: The target loss function is determined using at least one of the second disparity map and the third disparity map of the plurality of preset training images and the first disparity map of the plurality of preset training images.

28. The stereo matching method according to claim 27, wherein: The second disparity map is obtained by performing cost aggregation processing and disparity regression processing on the target cost volume in the stereo matching network.

29. The stereo matching method according to claim 27, wherein: The third disparity map is obtained by upsampling the initial cost volume in the stereo matching network and performing disparity regression processing.

30. The stereo matching method according to claim 27, wherein: The determining the target loss function by using at least one of the second disparity map and the third disparity map of the plurality of preset training images and the first disparity map of the plurality of preset training images includes: The target loss function is obtained by performing weighted summation using the loss function of at least one of the second disparity map and the third disparity map and the loss functions of the first disparity maps of multiple preset training images.

31. The stereo matching method according to claim 30, wherein: The weight of the loss function of the first disparity map of the plurality of preset training images is greater than the weight of the loss function of at least one of the second disparity map and the third disparity map.

32. A stereo matching device, characterized in that: The system comprises a processor, a memory, and a computer program or instruction, wherein the computer program or instruction is stored in the memory and executed by the processor to implement the stereo matching method according to any one of claims 1 to 31.

33. A computer-readable storage medium, characterized in that A computer program or instruction is stored thereon, and the computer program or instruction is executed by a processor to implement the stereo matching method according to any one of claims 1 to 31.

34. A computer program product, characterized in that The computer program or instructions are executed by a processor to implement the stereo matching method according to any one of claims 1 to 31.

35. A vehicle, characterized in that: Including the stereo matching device described in claim 32.