Image reconstruction method for automatic driving and related equipment

Through improved diffusion model and optimization matching cost method of area division, the problem of insufficient reconstruction accuracy of sparse views and weak texture areas is solved, the accuracy of three-dimensional reconstruction and the adaptability of complex scenes are improved, and more efficient image reconstruction effect is achieved.

CN120147520APending Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204486.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing multi-view stereo reconstruction technology has problems such as insufficient reconstruction accuracy and large local errors in sparse input views and weak texture areas, which is difficult to effectively improve the detail recovery ability of weak texture areas and the adaptability of complex scenes.

Method used

The improved diffusion model is used to complete the sparse views, generate dense view sequences, and optimize the matching cost through area division and multiple rounds of propagation, and three-dimensional reconstruction is carried out in combination with point cloud fusion algorithm. The improved diffusion model is used to generate a visually consistent dense view sequence, which solves the problem of information loss caused by sparse views, and improves the detail recovery ability of weak texture areas and the adaptability of complex scenes.

Benefits of technology

Through the improved diffusion model and area division optimization matching cost, the accuracy of three-dimensional reconstruction and the integrity of the reconstruction image are significantly improved, and the adaptability and computing efficiency to complex scenes are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147520A_ABST
    Figure CN120147520A_ABST
Patent Text Reader

Abstract

The invention provides an image reconstruction method for automatic driving and related equipment. The method comprises the following steps: acquiring a sparse view and camera pose information of a target automatic driving vehicle; complementing the sparse view based on the camera pose information by using the improved diffusion model to obtain a dense view sequence; randomly carrying out primary propagation on pixel points in the dense view sequence, carrying out region division on the dense view sequence by using depth information in a primary propagation result, and synchronously carrying out multi-round propagation by taking each pixel point in an obtained flat region, a boundary region and a shielding region as a propagation starting point, calculating the matching cost of each pixel point, optimizing the propagation result through the calculated matching cost of each pixel point, obtaining plane hypothesis data, fusing the plane hypothesis data through a point cloud fusion algorithm to obtain a three-dimensional point cloud model, and performing three-dimensional reconstruction on the three-dimensional point cloud model to obtain an image reconstruction result of the target autonomous vehicle; the detail recovery capability of a weak texture area is improved, local errors are reduced, and the adaptability of a complex scene is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image reconstruction, and particularly relates to an image reconstruction method and related devices for autonomous driving. Background Art

[0002] Multi-View Stereo (MVS) is one of the core tasks in the field of computer vision. Its main goal is to reconstruct the three-dimensional geometric structure of a scene from images taken from multiple viewpoints, simply referred to as 3D reconstruction. MVS technology has extensive application value in fields such as autonomous driving, virtual reality, digital content creation, and architectural engineering, and has played an important role especially in scenarios that require high-precision 3D models. However, in practical applications, MVS technology faces significant challenges in sparse input views and weak texture regions, resulting in limited reconstruction quality.

[0003] Traditional MVS methods are based on an efficient neighborhood matching technique of random search and local propagation. Its main processes include the following steps: randomly initializing plane hypotheses, propagation and update, matching cost calculation, iterative optimization, and point cloud fusion and reconstruction. This method has high computational efficiency and low memory requirements, and is particularly suitable for high-resolution images and large-scale scenes. However, it has obvious deficiencies in sparse input views and weak texture regions: in under-sampled areas, random initialization may not cover the global optimal solution; in weak texture regions, the matching cost function lacks significance and is prone to falling into local minima, resulting in a decrease in reconstruction accuracy.

[0004] There are also some studies that optimize the reconstruction accuracy in weak texture regions through methods such as geometric priors, multi-scale structures, and dynamic receptive field adjustment. The multi-scale structure can refine the depth estimation results layer by layer, enhance the receptive field coverage at low-resolution levels, and improve the detail reconstruction ability at high-resolution levels. At the same time, reliable pixel plane hypotheses are generated through geometric priors, providing guidance for depth estimation in weak texture regions and alleviating the matching ambiguity problem. In addition, the dynamic receptive field adjustment mechanism covers more reliable pixel points by adjusting the shape and range of patches and introduces pixel reliability classification, effectively enhancing the matching robustness. These improvements have, to a certain extent, improved the depth estimation quality and geometric consistency in weak texture regions and reduced the artifact phenomenon caused by insufficient texture.

[0005] Although the above methods have made significant progress in improving the reconstruction accuracy of weak texture regions, there are still certain limitations. Although the multi-scale structure can enhance the receptive field coverage of weak texture regions, the loss of detail information during the image scaling process is inevitable, affecting the three-dimensional reconstruction accuracy of high-resolution scenes. In addition, the dynamic receptive field adjustment is highly dependent on specific parameters and may not guarantee the best match in some special scenes, still requiring further optimization. Therefore, how to more comprehensively improve the detail recovery ability of weak texture regions, reduce local errors, and enhance the adaptability to complex scenes remains an important challenge in the current research of MVS technology. Summary of the Invention

[0006] The present invention provides an image reconstruction method and related devices for autonomous driving, aiming to improve the detail recovery ability of weak texture regions, reduce local errors, and enhance the adaptability to complex scenes.

[0007] To achieve the above object, the present invention provides an image reconstruction method for autonomous driving, including:

[0008] Step 1, obtaining the sparse view of the target autonomous driving vehicle and the camera pose information for capturing the sparse view;

[0009] Step 2, based on the camera pose information, using the improved diffusion model to complete the sparse view to obtain a dense view sequence;

[0010] Step 3, randomly propagating the pixel points in the dense view sequence to obtain the initial propagation result, and dividing the dense view sequence based on the depth information in the initial propagation result to obtain the flat region, boundary region, and occlusion region;

[0011] Step 4, synchronously performing multiple rounds of propagation with each pixel point in the flat region, boundary region, and occlusion region as the propagation starting point to obtain the propagation result, and optimizing the propagation result by calculating the matching cost of each pixel point to obtain the plane hypothesis data, where the plane hypothesis data includes the normal vector and depth;

[0012] Step 5, fusing the plane hypothesis data through a point cloud fusion algorithm to obtain a three-dimensional point cloud model, and performing three-dimensional reconstruction on the three-dimensional point cloud model to obtain the image reconstruction result of the target autonomous driving vehicle.

[0013] Furthermore, the improved diffusion model includes:

[0014] The first encoder, the second encoder, the U-net network, and the spatio-temporal decoder;

[0015] The input ends of the first encoder and the second encoder are both the input ends of the improved diffusion model;

[0016] The output end of the first encoder and the output end of the second encoder are both connected to the input end of the U-net network;

[0017] The output end of the U-net network is connected to the input end of the spatio-temporal decoder;

[0018] The output end of the spatio-temporal decoder is the output end of the improved diffusion model.

[0019] Furthermore, step 2 includes:

[0020] Input the sparse view and camera pose information into the first encoder for processing to obtain conditional input features, and input the conditional input features into the U-net network;

[0021] Input the sparse view into the second encoder for processing to obtain latent features, and input the latent features into the U-net network;

[0022] The U-net network diffuses and denoises the sparse view based on the conditional input features and latent features and then inputs it into the spatio-temporal decoder;

[0023] The spatio-temporal decoder decodes the diffused and denoised sparse view to obtain a dense view sequence.

[0024] Furthermore, the U-net network includes a first attention module, an intermediate module, and a second attention module connected in sequence;

[0025] The first input end of the first attention module is connected to the output end of the second encoder;

[0026] The second input end of the first attention module, the second input end of the intermediate module, and the second input end of the second attention module are all connected to the output end of the first encoder;

[0027] The output end of the second attention module is connected to the input end of the spatio-temporal decoder.

[0028] Furthermore, partitioning the dense view sequence according to the depth information in the initial propagation result includes:

[0029] According to the formula Calculate the depth gradient of each pixel point in the dense view sequence, where represents the depth gradient of the pixel point, represents the gradient of the pixel point in the x direction, represents the gradient of the pixel point in the y direction, and d(x, y) represents the depth value of the pixel point;

[0030] According to the formula Calculate the neighborhood texture change of each pixel point in the dense view sequence, where σ tIndicates the neighborhood texture change, I i Indicates the gray value of the i-th pixel point Indicates the average gray value of all pixels in the neighborhood where the i-th pixel point is located, and N represents the total number of pixel points in the neighborhood where the i-th pixel point is located;

[0031] Based on the depth gradient and neighborhood texture change, divide the flat areas and boundary areas in the dense view sequence;

[0032] According to the formula Calculate the projection position difference of the same pixel point in different perspectives in the dense view sequence to obtain the projection error, where e occl Indicates the projection error Indicates the position projection of the pixel point in the reference perspective Indicates the corresponding position where the pixel point in other perspective j is projected onto the pixel plane, and M represents the number of source perspectives participating in the multi-view consistency check;

[0033] Based on the projection error, divide the occlusion areas in the dense view sequence.

[0034] Furthermore, the matching cost calculated for each pixel point includes:

[0035] The expression for calculating the matching cost of each pixel point in the flat area is:

[0036]

[0037] The expression for calculating the matching cost of each pixel point in the boundary area is:

[0038]

[0039] The expression for calculating the matching cost of each pixel point in the occlusion area is:

[0040]

[0041] Among them, C NCC Indicates the matching cost of the pixel point in the flat area, J i Indicates the pixel value of the source perspective image at the i-th pixel point Indicates the average pixel value of each area in the source perspective image, C edge Indicates the matching cost of the pixel point in the boundary area, w i Indicates the weight coefficient, C occl Indicates the matching cost of the pixel point in the occlusion area, C i Indicates the matching cost of the i-th perspective, e i Indicates the projection error of the i-th pixel point, T occl Indicates the preset threshold.

[0042] Furthermore, taking each pixel point in the flat area, boundary area, and occlusion area as a propagation starting point, multiple rounds of propagation are synchronously performed to obtain propagation results, including:

[0043] Taking each pixel point in the flat area, boundary area, and occlusion area as a propagation starting point;

[0044] Construct a circular area with a radius of R around the propagation starting point, and evenly divide the circular area into eight sector blocks;

[0045] Adopt an alternating search method from top to bottom and from bottom to top to select the pixel point closest to the propagation starting point in each sector block for multiple rounds of propagation to obtain propagation results.

[0046] Furthermore, the stop conditions for propagation are:

[0047] When the matching costs of each pixel point in the boundary area or occlusion area converge, stop propagation;

[0048] When exceeding the set propagation area, stop propagation;

[0049] When the depth gradient change between the propagation starting point and the selected pixel point exceeds a preset threshold, stop propagation.

[0050] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements an image reconstruction method for autonomous driving.

[0051] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements an image reconstruction method for autonomous driving.

[0052] The above solution of the present invention has the following beneficial effects:

[0053] The present invention obtains the sparse view of the target autonomous driving vehicle and the camera pose information for photographing the sparse view; based on the camera pose information, uses the improved diffusion model to complete the sparse view to obtain a dense view sequence; randomly propagates the pixel points in the dense view sequence to obtain the initial propagation result, and divides the dense view sequence based on the depth information in the initial propagation result to obtain a flat area, a boundary area, and an occlusion area; synchronously performs multiple rounds of propagation with each pixel point in the flat area, the boundary area, and the occlusion area as the propagation starting point to obtain the propagation result, and optimizes the propagation result by calculating the matching cost of each pixel point to obtain plane hypothesis data, where the plane hypothesis data includes a normal vector and a depth; fuses the plane hypothesis data through a point cloud fusion algorithm to obtain a three-dimensional point cloud model, and performs three-dimensional reconstruction on the three-dimensional point cloud model to obtain the image reconstruction result of the target autonomous driving vehicle; compared with the prior art, the present invention completes the sparse view through the improved diffusion model to generate a visually consistent dense view sequence, solves the problem of information loss caused by the inability of traditional methods to process sparse views, improves the detail recovery ability of weak texture areas, reduces local errors and enhances the adaptability to complex scenes, thereby improving the accuracy of three-dimensional reconstruction and the integrity of the reconstructed image.

[0054] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flowchart of an embodiment of the present invention;

[0056] Figure 2 It is a schematic structural diagram of the improved diffusion model in an embodiment of the present invention;

[0057] Figure 3 It is a schematic structural diagram of the terminal device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.

[0060] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", "connection" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0061] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0062] The present invention aims at the existing problems and provides an image reconstruction method and related devices for autonomous driving.

[0063] As Figure 1 shown, an embodiment of the present invention provides an image reconstruction method for autonomous driving, including:

[0064] Step 1, obtaining the sparse view of the target autonomous driving vehicle and the camera pose information for capturing the sparse view;

[0065] Step 2, based on the camera pose information, using the improved diffusion model to complete the sparse view to obtain a dense view sequence;

[0066] Step 3, randomly propagating the pixel points in the dense view sequence to obtain the initial propagation result, and dividing the dense view sequence into a flat area, a boundary area, and an occlusion area based on the depth information in the initial propagation result;

[0067] Step 4, synchronously performing multiple rounds of propagation with each pixel point in the flat area, the boundary area, and the occlusion area as the propagation starting point to obtain the propagation result, and optimizing the propagation result by calculating the matching cost of each pixel point to obtain the plane hypothesis data, where the plane hypothesis data includes the normal vector and the depth;

[0068] Step 5: Fuse the plane hypothesis data through a point cloud fusion algorithm to obtain a 3D point cloud model, and perform 3D reconstruction on the 3D point cloud model to obtain the image reconstruction result of the target autonomous driving vehicle.

[0069] Specifically, the sparse view of the target autonomous driving vehicle consists of images taken from different angles. The spatial position of each image can be determined through camera pose information, and the determination of camera pose can be obtained through common camera calibration methods or automatically estimated based on existing image processing techniques.

[0070] It should be noted that the function of Step 1 is to ensure that the geometric relationship between the images in the sparse view is as accurate as possible.

[0071] Since when the input images are sparse views, due to insufficient sampling area information, the reconstruction effect will be significantly degraded, usually manifested as problems such as matching failure, artifact generation, and inaccurate depth estimation, seriously affecting the accuracy and integrity of the 3D model; especially when reconstructing the input views, the depth maps generated from different perspectives may be different, which will cause the finally generated 3D point cloud model to be visually discontinuous. To address this problem, the embodiments of the present invention complement the sparse views by improving the diffusion model to generate high-resolution dense views with consistency, providing more complete and reliable image data for subsequent 3D reconstruction, and significantly improving the 3D reconstruction performance under sparse view input. Figure 1 Specifically, as shown, the improved diffusion model includes:

[0072] Specifically, as Figure 2 shown, the improved diffusion model includes:

[0073] A first encoder, a second encoder, a U-net network, and a spatio-temporal decoder;

[0074] The input ends of the first encoder and the second encoder are both the input ends of the improved diffusion model;

[0075] The output ends of the first encoder and the second encoder are both connected to the input end of the U-net network;

[0076] The output end of the U-net network is connected to the input end of the spatio-temporal decoder;

[0077] The output end of the spatio-temporal decoder is the output end of the improved diffusion model.

[0078] Specifically, Step 2 includes:

[0079] Input the sparse view and camera pose information into the first encoder for processing to obtain conditional input features, and input the conditional input features into the U-net network;

[0080] The sparse view is input into the second encoder for processing to obtain latent features, and the latent features are input into the U-net network;

[0081] The U-net network diffuses and denoises the sparse view based on the conditional input features and latent features and then inputs it into the spatio-temporal decoder;

[0082] The spatio-temporal decoder decodes the diffused and denoised sparse view to obtain a dense view sequence.

[0083] In the embodiment of the present invention, before the sparse view and camera pose information are input into the first encoder for processing, it further includes:

[0084] Define the sparse view as {I ref1 , I ref2 , …, I refi}, and the corresponding camera poses are {p ref1 , p ref2 , …, p refi};

[0085] Uniformly sample the camera poses {p refi-1 , p refi} of every two reference views to obtain a series of camera poses represents the trajectory from the current reference view to the next view, and the dimension of the input data is v ∈ R (T+2)×i×3×H×W .

[0086] In the embodiment of the present invention, the first encoder consists of a CLIP encoder and a cross-attention module. The CLIP encoder is used to extract high-level semantic features of each image in the sparse view, that is, conditional input features. The cross-attention module inputs the extracted high-level semantic features into the U-net network; the second encoder is a VAE encoder, which is used to compress the input view into the latent space, extract low-dimensional latent features, and introduce the latent features into the diffusion model through the classifier-free guidance mechanism to enhance the consistency of color and texture in the generation process.

[0087] Specifically, the U-net network is used to capture the spatio-temporal relationship between frames in the video sequence to ensure the visual consistency of the generated images, including a first attention module, an intermediate module, and a second attention module connected in sequence;

[0088] The first input end of the first attention module is connected to the output end of the second encoder;

[0089] The second input end of the first attention module, the second input end of the intermediate module, and the second input end of the second attention module are all connected to the output end of the first encoder;

[0090] The output end of the second attention module is connected to the input end of the spatio-temporal decoder.

[0091] In an embodiment of the present invention, both the first attention module and the second attention module are cross-frame spatio-temporal attention modules. The cross-frame spatio-temporal attention module can capture the continuity and consistency in the spatio-temporal dimension of the video frame sequence. The input is the feature representation of each frame extracted from the video sequence, including spatial and temporal information related to adjacent frames. Then, by calculating the relationship between each frame and other frames, it ensures the information transmission and consistency between adjacent frames, and considers the spatio-temporal dependence between them. Finally, it outputs a coherent video frame sequence based on the spatio-temporal dependence relationship. The intermediate module is a 3D residual convolution module, which is used to process in the spatial and temporal dimensions, can better capture the dynamic characteristics of the data, extract features simultaneously in the three dimensions of image and time, and then add residual connections in the network to solve the problem of gradient disappearance and finally generate a more accurate spatio-temporal feature map.

[0092] Before processing the sparse view in the embodiment of the present invention, it is necessary to train the improved diffusion model. The training dataset needs to provide a rich high-resolution image sequence for the model through multi-view rendering. At the same time, the diversity and coverage of the data are ensured by randomly setting the elevation angle and uniformly distributing the azimuth angle, laying a foundation for subsequent image generation. The training objective is to learn the denoising process through the improved diffusion model, and the loss function is:

[0093]

[0094] where, z t = α t z + σ t ∈ *, z represents the true latent variable, ∈ ∈ N(0, I), α t and σ t define the noise corresponding to the step size t.

[0095] Since directly decoding the output image sequence of the improved diffusion model with a VAE decoder will result in problems such as time inconsistency, blurring, and color deviation, it is necessary to introduce a spatio-temporal decoder to convert the processed feature representation back into a high-quality output image sequence and ensure the continuity in the time (between frames) and space (local consistency of the image) dimensions.

[0096] Specifically, the spatio-temporal decoder includes a temporal convolution layer and a color normalization layer. The temporal convolution layer ensures a smooth transition between frames in the generated image sequence, and the color normalization layer uniformly adjusts the colors of the generated images to ensure image color consistency and high-quality output.

[0097] Most preferably, the dense view sequence is regionally divided based on the depth information in the initial propagation result, including:

[0098] According to the formula Calculate the depth gradient of each pixel in the dense view sequence. The depth gradient describes the degree of change in the depth value in the image by calculating the depth values of adjacent pixels, and is usually used to measure the changes on the object surface or boundary. Among them, represents the depth gradient of the pixel, represents the gradient of the pixel in the x direction, represents the gradient of the pixel in the y direction, d(x, y) represents the depth value of the pixel;

[0099] According to the formula Calculate the neighborhood texture change of each pixel in the dense view sequence. The neighborhood texture change refers to the degree of change in the pixel gray value within the neighborhood of a certain pixel, and measures the texture features of the image through the standard deviation of the pixel gray value. It can reflect the complexity and local changes of the texture, and is used to judge whether the area is a region with weak or strong texture. Among them, σ t represents the neighborhood texture change, I i represents the gray value of the i-th pixel, represents the average gray value of all pixels within the neighborhood where the i-th pixel is located, N represents the total number of pixels within the neighborhood where the i-th pixel is located;

[0100] Divide the flat areas and boundary areas in the dense view sequence based on the depth gradient and neighborhood texture change;

[0101] Since part of the object is blocked by other objects, it is impossible to correctly obtain the depth information of this area. In order to identify the occlusion area, according to the formula Calculate the projection position difference of the same pixel in different perspectives in the dense view sequence to obtain the projection error. Among them, e occl represents the projection error, represents the position projection of the pixel in the reference perspective, represents the corresponding position where the pixel in other perspective j is projected onto the pixel plane, and M represents the number of source perspectives participating in the multi-view consistency check;

[0102] Divide the occlusion areas in the dense view sequence based on the projection error.

[0103] Specifically, the neighborhood texture variation represents the degree of dispersion of the gray-scale distribution within the region. The larger the standard deviation, the stronger the texture variation in the region. A smaller standard deviation indicates weaker texture in the region, which may be a flat area. The depth gradient measures the rate of change of depth in the image and is used to distinguish the boundaries of objects and relatively flat areas. Combining the two as an evaluation index helps to effectively distinguish different regions in 3D reconstruction, thereby providing effective guidance for matching cost optimization, region adaption, and pixel propagation, etc. The projection error reflects the difference in the depth estimation results of a certain pixel point under different viewpoints. In summary, by dividing the region based on the depth gradient and neighborhood texture variation, when the depth gradient and the neighborhood texture variation σ t ≤T t , the region where the pixel point is located is a flat area; when the depth gradient and the neighborhood texture variation σ t ≥T t , the region where the pixel point is located is a boundary area; when the projection error e occl >T occl , the region where the pixel point is located is an occluded area, where T g , T t , T occl are all preset thresholds.

[0104] Since the geometric characteristics and matching requirements of different regions vary significantly, the traditional unified matching cost calculation method is difficult to fully meet the requirements of each region, easily leading to a decrease in matching accuracy or limited computational efficiency. The embodiments of the present invention achieve dynamic optimization of the cost calculation method by dividing regions, enabling it to focus on efficiency for flat areas, enhance the ability to capture geometric features for boundary areas, and improve multi-view consistency for occluded areas. Ultimately, a matching cost optimization mechanism with both robustness and adaptability is constructed. This connection design effectively improves the accuracy, integrity of depth estimation and the running efficiency of the algorithm in 3D reconstruction.

[0105] Specifically, the matching cost calculated for each pixel point includes:

[0106] Since flat areas usually have a small depth gradient and low texture variation, and the gray-scale value differences between pixels are not significant, the matching information is relatively scarce. In such areas, complex matching cost calculation methods cannot significantly improve the matching effect but will increase the computational burden. Therefore, the normalized cross-correlation (NCC) is used as the matching cost calculation formula. By means of a simple and efficient similarity measurement method, while maintaining high computational efficiency, it ensures basic matching quality, effectively avoiding over-computation and providing support for efficient depth estimation in large-scale scenes. The expression for calculating the matching cost of each pixel point in the flat area is:

[0107]

[0108] Due to the large depth gradient in the boundary region, which has significant directional characteristics and is an important part of the object's contour and geometric structure; however, due to the drastic changes in depth and texture, traditional patch matching methods with fixed shapes cannot effectively capture boundary details, easily leading to edge blurring or mismatching; for this reason, a direction-weighted NCC matching cost calculation formula is introduced. By dynamically adjusting the weights to enhance the sensitivity to the gradient direction, the matching cost can better reflect the geometric characteristics of the boundary region, significantly improving the accuracy of boundary matching and ensuring clear edges and rich details in the 3D reconstruction results. The matching cost expression for each pixel point in the boundary region is calculated as follows:

[0109]

[0110] In the multi-view matching of occluded regions, depth inconsistencies often occur. Affected by noise and occluders, it is easy to cause instability or even mismatching of the matching cost; to address this problem, multi-view consistency checking is used to calculate the projection error, eliminating the view matching values with large errors and only retaining the matching costs with high consistency; through this optimization, noise interference can be effectively reduced, the influence of incorrect depth values can be minimized, thereby significantly improving the reconstruction quality of occluded regions and enhancing the robustness and accuracy of the multi-view reconstruction algorithm. The matching cost expression for each pixel point in the occluded region is calculated as follows:

[0111]

[0112] Among them, C NCC represents the matching cost of pixel points in the flat region, J i represents the pixel value of the source view image at the i-th pixel point, represents the average pixel value of each region in the source view image, C edge represents the matching cost of pixel points in the boundary region, w i represents the weight coefficient, C occl represents the matching cost of pixel points in the occluded region, C i represents the matching cost of the i-th view, e i represents the projection error of the i-th pixel point, T occl represents the preset threshold.

[0113] Most preferably, each pixel point in the flat region, boundary region, and occluded region is used as the propagation starting point to synchronously perform multiple rounds of propagation to obtain the propagation results, including:

[0114] Using each pixel point in the flat region, boundary region, and occluded region as the propagation starting point;

[0115] Construct a circular area with a radius of R around the propagation starting point, and evenly divide the circular area into eight sector blocks;

[0116] Adopt an alternating search method from top to bottom and from bottom to top to select the pixel point closest to the propagation starting point in each sector block for multiple rounds of propagation to obtain the propagation result.

[0117] Specifically, an improved search strategy is adopted during the initial propagation to improve the propagation efficiency and reduce the computational complexity; first, with the propagation starting point as the center, construct a circular area with a search radius of R, and evenly divide the circular area into eight sector blocks, which are respectively labeled as eight regions: up, upper right, right, lower right, down, lower left, left, and upper left. Start the search from the pixel closest to the propagation starting point and located above in each region, and perform the propagation in the order from top to bottom to preferentially explore the closer regions;

[0118] Subsequently, switch to the second row of pixel points in the sector area and adopt a search order from bottom to top; through this up-and-down alternating search method, the search area can be effectively covered and the propagation path can be optimized, reducing unnecessary calculations;

[0119] After the traversal is completed, compare to obtain the pixel point with the minimum matching cost. If the matching cost value of this point is less than that of the propagation starting point, then use it as the plane hypothesis data, and the plane hypothesis data includes the normal vector and depth. Otherwise, it is considered that there are no valid propagation points in this region.

[0120] It should be noted that if the search circle cannot be constructed around the propagation starting point, it is also considered that there are no valid propagation points in this region.

[0121] During this propagation, only the matching cost is used for judgment, without calculating the depth gradient change, reducing the computational complexity. And through the ordered search order and dynamic adjustment, the computational burden in the initial propagation is reduced, while ensuring the comprehensiveness of the search coverage range and the accuracy of cost optimization, thereby improving the propagation efficiency and accuracy.

[0122] In the embodiments of the present invention, not all pixel points need to be propagated for the second time, because only when the matching does not converge or there is uncertainty, it is necessary to further optimize the propagation result to improve the accuracy and efficiency. Therefore, the propagation judgment mechanism introduced in the embodiments of the present invention is:

[0123] For each region, judge one by one whether there are valid propagation points;

[0124] If there are no valid propagation points in a certain region, it is considered that this region needs to continue the propagation, and the propagation starts from the midpoint of the outermost periphery of the sector area;

[0125] For the area containing valid propagation points, the candidate points obtained from the initial propagation are selected as the propagation starting points. Then, a straight line is connected between the propagation starting point and the selected pixel points, and a zigzag propagation is formed along this line with a vertical offset, so that the search range can be expanded, covering a wider area, increasing the matching accuracy. The propagation area is set to be centered on the propagation starting point, and a propagation radius R is defined around it. p The propagation is only carried out within this area. When the area is exceeded, the propagation stops, and it is marked that there are no valid propagation points in the current area. Further, it is judged whether the matching cost exceeds the set threshold. When the matching cost is greater than the threshold, it indicates that there may be uncertainties or errors in the current matching result, and further optimization is required through continuous propagation. Especially in the boundary or occlusion areas, when points with matching costs meeting the conditions are found during the propagation process, the propagation will terminate immediately, and they are marked as valid propagation points, thus reducing unnecessary calculation operations and improving the calculation efficiency.

[0126] The matching cost condition is as follows:

[0127] cost < T c and cost < cost i

[0128] where cost is the matching cost of the pixel point where the propagation reaches, cost i is the matching cost value of the propagation starting point, and T c is the cost threshold.

[0129] This optimization refines the depth estimation by introducing more neighborhood information, improves the matching accuracy, and restores the details of complex areas, thereby reducing mis-matches and enhancing the accuracy of depth estimation.

[0130] This propagation judgment mechanism can also effectively reduce unnecessary computational overhead, and only propagates in the areas where the matching cost has not converged, avoiding redundant propagation operations. Especially in the converged areas, it can significantly improve the computational efficiency and optimize the propagation process. This strategy not only ensures the matching accuracy but also improves the computational efficiency, ensuring the fast response and efficient execution of the algorithm in complex scenarios.

[0131] In summary, the stop conditions for propagation are as follows:

[0132] After the initial propagation, when the matching costs of each pixel point in the boundary area or occlusion area converge, the propagation stops;

[0133] When the set propagation area is exceeded, the propagation stops. This propagation area is a circular area with a radius of 2R constructed around the propagation starting point;

[0134] During the propagation process, when the depth gradient change corresponding to the propagation starting point and the selected pixel point exceeds a preset threshold, the propagation stops to avoid ineffective propagation in the boundary or occlusion area.

[0135] Specifically, in step 5, the plane hypothesis data is fused through a point cloud fusion algorithm to obtain a three-dimensional point cloud model. The point cloud fusion algorithm used is a common fusion method in the field of visual fusion. The embodiments of the present invention do not involve improving the steps of the algorithm, so the fusion process will not be elaborated one by one, and the three-dimensional point cloud model is three-dimensionally reconstructed to obtain the image reconstruction result of the target autonomous vehicle. The image reconstruction result is a three-dimensional image of the target autonomous vehicle.

[0136] In the embodiments of the present invention, the sparse view of the target autonomous vehicle and the camera pose information for capturing the sparse view are obtained; based on the camera pose information, the improved diffusion model is used to complete the sparse view to obtain a dense view sequence; the pixel points in the dense view sequence are randomly propagated to obtain the initial propagation result, and the dense view sequence is regionally divided based on the depth information in the initial propagation result to obtain the flat area, boundary area, and occlusion area; each pixel point in the flat area, boundary area, and occlusion area is used as the propagation starting point to synchronously perform multiple rounds of propagation to obtain the propagation result, and the propagation result is optimized by calculating the matching cost of each pixel point to obtain the plane hypothesis data, which includes the normal vector and depth; the plane hypothesis data is fused through a point cloud fusion algorithm to obtain a three-dimensional point cloud model, and the three-dimensional point cloud model is three-dimensionally reconstructed to obtain the image reconstruction result of the target autonomous vehicle; compared with the prior art, in the embodiments of the present invention, the improved diffusion model is used to complete the sparse view to generate a visually consistent dense view sequence, solving the problem of information loss caused by the inability of traditional methods to process sparse views, improving the detail recovery ability in weakly textured areas, reducing local errors and enhancing the adaptability to complex scenes, thereby improving the accuracy of three-dimensional reconstruction and the integrity of the reconstructed image.

[0137] The embodiments of the present invention also provide a terminal device, as Figure 3 shown. The terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 only one processor is shown), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the above-mentioned image reconstruction method for autonomous driving is implemented.

[0138] The terminal device D10 may be a computing device such as a desktop computer, a notebook, a palm computer, a server, a server cluster, and a cloud server. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art can understand that Figure 3 merely examples of the terminal device D10, which do not constitute a limitation on the terminal device D10, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0139] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and the processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application-specific integrated circuits (ASIC, Application Specific Integrated Circuit), off-the-shelf programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0140] The memory D101 may be an internal storage unit of the terminal device D10 in some embodiments, such as the hard disk or memory of the terminal device D10. The memory D101 may also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC, Smart Media Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0141] It should be noted that the content such as information interaction and execution process between the above-mentioned devices / units, due to being based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.

[0142] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0143] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements an image reconstruction method for autonomous driving.

[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0145] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image reconstruction method for autonomous driving, characterized in that: include: Step 1, obtaining a sparse view of a target autonomous driving vehicle and camera pose information for capturing the sparse view; Step 2, based on the camera pose information, using the improved diffusion model to complete the sparse view to obtain a dense view sequence; Step 3, randomly propagating the pixels in the dense view sequence to obtain a primary propagation result, and dividing the dense view sequence into regions according to the depth information in the primary propagation result to obtain a flat region, a boundary region, and an occluded region; Step 4, taking each pixel point in the flat area, the boundary area and the occluded area as a propagation starting point, synchronously performing multiple rounds of propagation to obtain a propagation result, and optimizing the propagation result by calculating the matching cost of each pixel point to obtain plane hypothesis data, wherein the plane hypothesis data includes a normal vector and a depth; Step 5: Fusing the plane hypothesis data through a point cloud fusion algorithm to obtain a three-dimensional point cloud model, and reconstructing the three-dimensional point cloud model to obtain an image reconstruction result of the target autonomous driving vehicle.

2. The image reconstruction method for autonomous driving according to claim 1, characterized in that: The improved diffusion model includes: First encoder, second encoder, U-net network, spatiotemporal decoder; The input end of the first encoder and the input end of the second encoder are both input ends of the improved diffusion model; The output end of the first encoder and the output end of the second encoder are both connected to the input end of the U-net network; The output end of the U-net network is connected to the input end of the spatiotemporal decoder; The output end of the spatiotemporal decoder is the output end of the improved diffusion model.

3. The image reconstruction method for autonomous driving according to claim 2, characterized in that: The step 2 comprises: Inputting the sparse view and the camera pose information into the first encoder for processing to obtain conditional input features, and inputting the conditional input features into the U-net network; Inputting the sparse view into the second encoder for processing to obtain potential features, and inputting the potential features into the U-net network; The U-net network diffuses and denoises the sparse view based on the conditional input features and the potential features, and then inputs the diffused view into the spatiotemporal decoder; The spatiotemporal decoder decodes the sparse view after diffusion denoising to obtain a dense view sequence.

4. The image reconstruction method for autonomous driving according to claim 3, characterized in that: The U-net network includes a first attention module, an intermediate module, and a second attention module connected in sequence; A first input terminal of the first attention module is connected to an output terminal of the second encoder; The second input end of the first attention module, the second input end of the intermediate module, and the second input end of the second attention module are all connected to the output end of the first encoder; An output of the second attention module is connected to an input of the spatiotemporal decoder.

5. The image reconstruction method for autonomous driving according to claim 4, characterized in that: The dense view sequence is divided into regions according to the depth information in the initial propagation result, including: According to the formula Calculate the depth gradient of each pixel in the dense view sequence, where: Represents the depth gradient of the pixel, Represents the gradient of the pixel in the x direction, Represents the gradient of the pixel in the y direction, and d(x,y) represents the depth value of the pixel; According to the formula Calculate the neighborhood texture change of each pixel in the dense view sequence, where σ t Indicates the neighborhood texture change, I i represents the gray value of the i-th pixel, represents the grayscale mean of all pixels in the neighborhood where the i-th pixel is located, and N represents the total number of pixels in the neighborhood where the i-th pixel is located; Dividing a flat area and a boundary area in the dense view sequence based on the depth gradient and the neighborhood texture change; According to the formula Calculate the difference in projection position of the same pixel in the dense view sequence at different viewing angles to obtain the projection error, where e occl represents the projection error, Represents the position projection of the pixel point in the reference perspective, represents the corresponding position of the pixel point in other view j projected onto the pixel plane, and M represents the number of source views involved in the multi-view consistency check; An occlusion area in the dense view sequence is divided based on the projection error.

6. The image reconstruction method for autonomous driving according to claim 5, characterized in that: The matching cost of each pixel is calculated, including: The matching cost expression for each pixel in the flat area is calculated as follows: The matching cost expression for each pixel in the boundary area is calculated as follows: The matching cost expression for each pixel in the occluded area is calculated as follows: Among them, C NCC represents the matching cost of pixels in the flat area, J i Represents the pixel value of the source perspective image at the i-th pixel point, Represents the average pixel value of each region in the source view image, C edge represents the matching cost of pixels in the boundary area, w i represents the weight coefficient, C occl represents the matching cost of pixels in the occluded area, C i represents the matching cost of the i-th view, e i represents the projection error of the i-th pixel, T occl Indicates the preset threshold.

7. The image reconstruction method for autonomous driving according to claim 6, characterized in that: Multiple rounds of propagation are synchronously performed with each pixel point in the flat area, the boundary area, and the occluded area as the propagation starting point to obtain propagation results, including: Taking each pixel point in the flat area, the boundary area and the occluded area as a propagation starting point; Construct a circular area with a radius of R around the propagation starting point, and divide the circular area into eight sector blocks; The pixel points closest to the propagation starting point are selected in each sector block in an alternating search mode from top to bottom and from bottom to top for multiple rounds of propagation to obtain the propagation result.

8. The image reconstruction method for autonomous driving according to claim 7, characterized in that: The stopping conditions for propagation are: When the matching cost of each pixel point in the boundary area or the occluded area converges, the propagation is stopped; When the set propagation area is exceeded, the propagation will be stopped; When the depth gradient change between the propagation starting point and the selected pixel point exceeds a preset threshold, the propagation is stopped.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image reconstruction method for autonomous driving is implemented as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image reconstruction method for autonomous driving is implemented as described in any one of claims 1 to 8.