Three-dimensional reconstruction method, device and equipment for real-time environment of vehicle

By acquiring multi-view image frame sequences of the vehicle's surrounding environment for 3D reconstruction and rendering fusion, the real-time performance and high fidelity issues of existing 3D reconstruction technologies are solved. This achieves the combination of realistic static backgrounds and dynamic objects, improving user experience and the clarity of information presentation.

CN121962464APending Publication Date: 2026-05-01XG TECHNOLOGIES PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XG TECHNOLOGIES PTE LTD
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the 3D reconstruction of the vehicle's surrounding environment lacks real-time performance and high fidelity, resulting in a poor user experience, especially when rendering symbolic models on a virtual black/gray background, where there is a lack of immersion and scene context information.

Method used

By acquiring target image frame sequences from different shooting angles of the vehicle's current surrounding environment, 3D reconstruction is performed based on deep learning and rendering technology. Combining real static background with dynamic object representation, a 3D scene containing a real static background is generated and rendered and fused to form a target image corresponding to the current surrounding environment.

Benefits of technology

It enables real-time, high-fidelity 3D scene visualization of the vehicle's surroundings, enhancing the user experience, maintaining the immersiveness and clarity of information in a realistic driving environment, and avoiding information overload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962464A_ABST
    Figure CN121962464A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional reconstruction method, device and equipment for a real-time environment of a vehicle. The method comprises the following steps: acquiring a target image frame sequence of a current surrounding environment of the vehicle; performing three-dimensional reconstruction based on the target image frame sequence to obtain a three-dimensional scene representation of the current surrounding environment; wherein the three-dimensional scene representation comprises a static background representation and a dynamic object representation; and rendering the static background representation and the target dynamic representation to obtain a target picture corresponding to the current surrounding environment. According to the scheme, the three-dimensional scene containing the real static background is reconstructed based on the image frame sequence of the real surrounding environment, so that the static scene context of the real driving environment is completely reserved in the picture construction process, the real static background is combined with the dynamic elements, and the real driving environment is constructed. The problem of immersion deficiency caused by rendering a symbolized model on a virtual black / gray background in related technologies is avoided, and user experience can return to a real driving environment through static scene context information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer vision and vehicle-mounted human-machine interaction, specifically to a method, apparatus, and device for three-dimensional reconstruction of a vehicle's real-time environment. Background Technology

[0002] Advanced Driver Assistance Systems (ADAS) and autonomous driving systems are typically equipped with Surrounding Reality (SR) functionality, which displays the driver's understanding of the vehicle's surroundings on the in-vehicle Human-Machine Interface (HMI) screen. However, achieving real-time, high-fidelity 3D reconstruction of the vehicle's surroundings remains a critical challenge. Summary of the Invention

[0003] To address the aforementioned technical issues, this disclosure provides a method, apparatus, and device for real-time 3D reconstruction of a vehicle's surrounding environment, enabling real-time, high-fidelity 3D reconstruction of the vehicle's surrounding environment.

[0004] The first aspect of this disclosure provides a method for three-dimensional reconstruction of a vehicle's real-time environment, comprising: Acquire a sequence of target image frames of the vehicle's current surrounding environment; wherein, the sequence of target image frames includes image frames from different shooting angles; Based on the target image frame sequence, a 3D reconstruction is performed to obtain a 3D scene representation of the current surrounding environment; wherein, the 3D scene representation includes a static background representation and a dynamic object representation; The static background representation and the dynamic target representation are rendered to obtain a target image corresponding to the current surrounding environment; wherein, the dynamic target representation is the dynamic object representation or a preset standard 3D model.

[0005] A second aspect of this disclosure provides a three-dimensional reconstruction apparatus for a vehicle's real-time environment, comprising: The first acquisition module is used to acquire a sequence of target image frames of the vehicle's current surrounding environment; wherein, the sequence of target image frames includes image frames from different shooting angles; A 3D reconstruction module is used to perform 3D reconstruction based on the target image frame sequence to obtain a 3D scene representation of the current surrounding environment; wherein, the 3D scene representation includes a static background representation and a dynamic object representation; The image rendering module is used to render the static background representation and the target dynamic representation to obtain a target image corresponding to the current surrounding environment; wherein, the target dynamic representation is the dynamic object representation or a preset standard three-dimensional model.

[0006] A third aspect of this disclosure provides a computer-readable storage medium storing a computer program for performing the three-dimensional reconstruction method for the real-time environment of a vehicle as described in the first aspect.

[0007] A fourth aspect of this disclosure provides an electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the three-dimensional reconstruction method for the real-time environment of a vehicle as described in the first aspect above.

[0008] A fifth aspect of this disclosure provides a computer program product that, when executed by an instruction processor, performs a three-dimensional reconstruction method for the real-time environment of a vehicle as proposed in the first aspect of this disclosure.

[0009] The technical solution provided in this disclosure acquires a sequence of target image frames from different shooting angles of the vehicle's current surrounding environment, and performs 3D reconstruction based on this sequence to obtain a 3D scene representation including static background and dynamic object representations. Then, the static background representation and the target dynamic representation (dynamic object representation or a preset standard 3D model) are rendered to obtain a target image containing the vehicle and corresponding to the current surrounding environment. Since this solution reconstructs a 3D scene containing a real static background based on a sequence of image frames from the real surrounding environment, the static scene context of the real driving environment is completely preserved during the image construction process. This combines the real static background with dynamic elements, avoiding the immersive loss problem caused by rendering symbolic models on a virtual black / gray background in related technologies. It allows the user experience to return to a real driving environment through static scene context information. Thus, real-time, high-fidelity 3D scene visualization of the vehicle's surrounding environment is achieved, effectively improving the user experience. Attached Figure Description

[0010] Figure 1 This is an architectural flowchart of an augmented reality fusion display system based on forward reasoning reconstruction and real-time segmentation, provided by an exemplary embodiment of this disclosure.

[0011] Figure 2 This is a flowchart illustrating a three-dimensional reconstruction method for a vehicle's real-time environment provided in an exemplary embodiment of this disclosure.

[0012] Figure 3This is a flowchart illustrating a method for three-dimensional reconstruction of a vehicle's real-time environment provided in another exemplary embodiment of this disclosure.

[0013] Figure 4 This is a flowchart illustrating a method for three-dimensional reconstruction of a vehicle's real-time environment provided in another exemplary embodiment of this disclosure.

[0014] Figure 5 This is a flowchart illustrating a method for three-dimensional reconstruction of a vehicle's real-time environment provided in another exemplary embodiment of this disclosure.

[0015] Figure 6 This is a flowchart illustrating a method for three-dimensional reconstruction of a vehicle's real-time environment provided in another exemplary embodiment of this disclosure.

[0016] Figure 7 A schematic diagram of the structure of a three-dimensional reconstruction apparatus for a real-time vehicle environment provided as an exemplary embodiment of this disclosure.

[0017] Figure 8 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0018] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0019] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0020] Application Overview In related technologies, SR systems render and display symbolic, cartoonish models of vehicles, pedestrians, and lane lines against a virtual black or gray background. However, this method is completely detached from the real driving environment in which the vehicle is located, lacking immersion and sufficient scene context information, resulting in a poor user experience. Therefore, how to achieve real-time, high-fidelity 3D reconstruction of the vehicle's surrounding environment has become an urgent problem to be solved.

[0021] To address the aforementioned issues, the technical solution provided in this disclosure acquires a sequence of target image frames from different shooting angles of the vehicle's current surrounding environment. Based on this sequence, it performs 3D reconstruction to obtain a 3D scene representation including static background and dynamic object representations. Then, it renders the static background representation and the target dynamic representation (dynamic object representation or a preset standard 3D model) to obtain a target image containing the vehicle and corresponding to the current surrounding environment. Since this solution reconstructs a 3D scene containing a real static background based on a sequence of image frames from the real surrounding environment, the static scene context of the real driving environment is fully preserved during the image construction process. This combines the real static background with dynamic elements, avoiding the immersive loss problem caused by rendering symbolic models on a virtual black / gray background in related technologies. It allows the user experience to return to a real driving environment through static scene context information. Thus, it achieves real-time, high-fidelity 3D scene visualization of the vehicle's surrounding environment, effectively improving the user experience.

[0022] Exemplary System Figure 1 This is an architectural flowchart of an augmented reality fusion display system based on forward reasoning reconstruction and real-time segmentation, provided by an exemplary embodiment of this disclosure. Figure 1 As shown, a pipeline design enables the vehicle-side to complete the closed loop of "acquisition-reconstruction-segmentation-rendering fusion-display" in real time. The following is a brief explanation of the above architecture flowchart in four stages: Phase 1: Acquiring synchronized image sequences using onboard sensors. Specifically, a surround-view camera array is used to capture the vehicle's current surroundings, resulting in a synchronized initial image frame sequence.

[0023] The second stage involves processing the initial image frame sequence using the onboard computing unit. Specifically, the initial image frame sequence can be preprocessed to obtain a distortion-free target image frame sequence. Then, 3D reconstruction is performed based on the target image frame sequence to obtain a 3D scene representation of the current surrounding environment. Subsequently, the 3D scene representation can be segmented using a pre-defined deep learning network to obtain a first semantic label and a second semantic label for the 3D scene representation. Finally, a static background representation is separated from the 3D scene representation based on the first and second semantic labels.

[0024] The third stage involves rendering and blending using a rendering fusion engine. Specifically, a static background image is obtained by rendering a static background representation using a static scene renderer, and a symbolic dynamic foreground is obtained by rendering a standard 3D model based on real-time perceived dynamic object information and a symbolic SR renderer. Then, the static background image and the symbolic dynamic foreground are blended to obtain the target image, which is the final image to be displayed.

[0025] The fourth stage involves displaying the target image on the in-vehicle display screen. This creates an augmented reality driving view that includes real-world texture details while maintaining key dynamic information that is clear and easy to read.

[0026] It should be noted that the entire processing pipeline described above (reconstruction, segmentation, and rendering) is designed to be real-time and low-latency. For detailed descriptions of the above embodiments, please refer to the relevant explanations in the following method embodiments; these will not be repeated here.

[0027] The technical solution provided in this disclosure acquires a sequence of target image frames from different shooting angles of the vehicle's current surrounding environment, and performs 3D reconstruction based on this sequence to obtain a 3D scene representation including static background and dynamic object representations. Then, the separated static background representation and a preset standard 3D model are rendered to obtain a target image containing the vehicle image and corresponding to the current surrounding environment. Since this solution reconstructs a 3D scene containing a real static background based on a sequence of image frames from the real surrounding environment, the static scene context of the real driving environment is completely preserved during the image construction process. This combines the real static background with dynamic elements, avoiding the immersive loss problem caused by rendering symbolic models on a virtual black / gray background in related technologies. It allows the user experience to return to a real driving environment through static scene context information. Thus, real-time, high-fidelity 3D scene visualization of the vehicle's surrounding environment is achieved, effectively improving the user experience.

[0028] Moreover, this solution creatively combines realistically reconstructed static backgrounds with symbolic dynamic subjects, providing drivers with rich real-world environmental context while ensuring rapid identification of key dynamic information such as surrounding vehicles and pedestrians, thus avoiding information overload.

[0029] On the one hand, the real-time generated static 3D environment model is a high-value map data that can be used for crowdsourced high-definition map (HD Map) construction, driving behavior analysis, or the creation of digital twin cities; on the other hand, the dynamic part adopts lightweight symbolic rendering, which can effectively reduce the rendering load of the vehicle computing platform.

[0030] Exemplary methods Figure 2 This is a schematic flowchart illustrating a three-dimensional reconstruction method for a vehicle's real-time environment provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 2 As shown, it may include the following steps: Step 201: Obtain the target image frame sequence of the vehicle's current surrounding environment.

[0031] The target image frame sequence includes image frames from different shooting angles.

[0032] In some examples, a surround-view camera array can be used to synchronously acquire real-time images of the vehicle's surrounding environment to obtain an initial image frame sequence. This initial image frame sequence can then be preprocessed to obtain a target image frame sequence in real time, resulting in a high-quality target image frame sequence. For the preprocessing of the initial image frame sequence, please refer to the detailed description in the following embodiments; further elaboration is not required here.

[0033] • For example, the aforementioned surround-view camera group includes automotive-grade high dynamic range (HDR) cameras positioned at different locations on the vehicle in a surround-view layout to provide a 360-degree field of view without blind spots. For instance, six wide-angle cameras are respectively positioned at the front, left, right, top, left rear, and right rear of the vehicle.

[0034] Step 202: Perform 3D reconstruction based on the target image frame sequence to obtain a 3D scene representation of the current surrounding environment.

[0035] The 3D scene representation includes static background representation and dynamic object representation.

[0036] In some examples, a deep learning-based, forward-inference encoder-decoder architecture can be employed to perform 3D reconstruction based on a sequence of target image frames. For instance, this architecture could be a Visual Geometry Grounded Transformer (VGGT) or other possible neural network models. This architecture compresses the multi-view camera image stream from the input target image frame sequence into a compact latent representation describing the entire scene through an encoder network. Then, a decoder network directly transforms (decodes) this latent representation into the final 3D scene representation. Specific details can be found in the following embodiments, which will not be elaborated upon here.

[0037] For example, the aforementioned 3D scene representation can be a dense 3D point cloud and its attribute information, which may include color, normals, and reflectivity; as another example, the aforementioned 3D scene representation can be a discrete Gaussian sphere model and its attribute information, which may include color, shape, and volume. The above are merely exemplary descriptions of 3D scene representations, and the embodiments disclosed herein do not limit the scope of the representation.

[0038] Step 203: Render the static background representation and the dynamic target representation to obtain the target image corresponding to the current surrounding environment.

[0039] The target dynamic representation is either a dynamic object representation or a preset standard 3D model.

[0040] In some embodiments, when the target is dynamically represented as a dynamic object, the same rendering pipeline can be used to render both the static background representation and the target dynamic representation simultaneously to obtain a target image corresponding to the current surrounding environment; when the target is dynamically represented as a standard 3D model, different rendering pipelines are used to render the static background representation and the standard 3D model respectively, and the static background image obtained after rendering the static background representation and the symbolic dynamic foreground obtained after rendering the standard 3D model are fused to obtain the target image.

[0041] The technical solution provided in this disclosure acquires a sequence of target image frames from different shooting angles of the vehicle's current surrounding environment, and performs 3D reconstruction based on this sequence to obtain a 3D scene representation including static background and dynamic object representations. Then, the static background representation and the target dynamic representation (dynamic object representation or a preset standard 3D model) are rendered to obtain a target image containing the vehicle and corresponding to the current surrounding environment. Since this solution reconstructs a 3D scene containing a real static background based on a sequence of image frames from the real surrounding environment, the static scene context of the real driving environment is completely preserved during the image construction process. This combines the real static background with dynamic elements, avoiding the immersive loss problem caused by rendering symbolic models on a virtual black / gray background in related technologies. It allows the user experience to return to a real driving environment through static scene context information. Thus, real-time, high-fidelity 3D scene visualization of the vehicle's surrounding environment is achieved, effectively improving the user experience.

[0042] In some embodiments, such as Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 202 above may include the following steps: Step 2021: Encode the target image frame sequence based on the preset encoder network to obtain the target potential vector used to describe the current surrounding environment.

[0043] In some examples, the encoder network described above can be a Vision Transformer (ViT) or a Convolutional Neural Network (CNN). This encoder network, through a multi-layer self-attention mechanism and a feedforward neural network, deeply mines and fuses the local features and global semantic information of image frames from different perspectives within the target image frame sequence. For example, it effectively encodes static background features such as road textures, building outlines, and traffic signs, as well as dynamic object features such as vehicles and pedestrians. The final output is a target latent vector that accurately represents the three-dimensional structure and semantic information of the current surrounding environment. This target latent vector can be a fixed-dimensional latent vector, the dimension of which can be determined by the number of neurons in the encoder's output layer. For example, the target latent vector could be a 256-dimensional latent vector.

[0044] Step 2022: Obtain the historical latent vector of the previous moment.

[0045] In some examples, the aforementioned historical latent vectors may exist in a local cache or on a designated server. When performing 3D reconstruction based on continuous video streams, the historical latent vectors from the previous moment can be retrieved from the local cache or a designated server based on the timestamp. These historical latent vectors have the same dimension as the target latent vector at the current moment, and may include the 3D structure and semantic information of the surrounding environment from the previous moment.

[0046] Step 2023: Based on the preset decoder network, the target latent vector and the historical latent vector are forward propagated to obtain a 3D scene representation.

[0047] In some embodiments, for the target latent vector at the current moment, the target latent vector can be directly forward-propagated based on a pre-defined decoder network to obtain a 3D scene representation. However, for continuous video streams, to help the model better understand the dynamic changes in the scene and achieve tracking of dynamic objects and smooth scene evolution, the target latent vector and historical latent vectors can be used as inputs to the decoder network simultaneously. The decoder network fuses these two vectors, for example, through element-wise addition, concatenation, and transformation by a fully connected layer. The fused vector is then used as the initial state for forward propagation calculation through the decoder network. During forward propagation, the decoder network gradually decodes the abstract features contained in the fused vector into specific 3D geometric information and attribute information, thus obtaining the 3D scene representation. The 3D scene representation includes 3D geometric information and attribute information.

[0048] Based on the above embodiments, since high-dimensional image data can be compressed into low-dimensional target latent vectors through an encoder network, and then decoded by combining the target latent vectors and historical latent vectors to obtain an accurate 3D scene representation, it can not only reduce reconstruction errors caused by noise or occlusion in single-frame images, but also enhance the accuracy of dynamic object tracking through temporal correlation, making the 3D scene representation maintain a smooth transition in the time dimension, further improving the high fidelity of 3D reconstruction, and laying a solid foundation for rendering realistic and dynamically coherent target images while ensuring real-time performance. Moreover, the forward-inference-based reconstruction algorithm eliminates the dependence on offline optimization in traditional methods, enabling 3D reconstruction of dynamically changing scenes during vehicle movement, thus achieving true real-time in-vehicle reconstruction.

[0049] In some embodiments, when the target is dynamically represented as a standard three-dimensional model; such as Figure 4 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 203 above may include the following steps: Step 2031: Render the static background representation based on the first rendering pipeline to obtain a static background image.

[0050] For example, for a 3D engine on an embedded platform, the first rendering pipeline mentioned above can be Unity's High Definition Render Pipeline (HDRP) or Universal Render Pipeline (URP), or the mobile renderer of Unreal Engine; or a self-developed engine can be used for rendering.

[0051] In some examples, after separating the static background representation from the 3D scene representation, the static background representation can be rendered. When the static background representation is a Gaussian sphere model, it can be directly imported into the 3D engine associated with the first rendering pipeline. The engine's built-in rendering functions, such as lighting calculation, material assignment, and texture mapping, can be used to render the model and generate a realistic static background image. When a static background is represented as 3D point cloud data, specific point cloud rendering techniques can be employed. For example, shader-based sprite rendering assigns a sprite texture with transparency and size attributes to each point, allowing for rapid rendering of large amounts of point cloud data into a continuous visual effect under parallel GPU processing, thus obtaining a static background image. Alternatively, splatting rendering can be used, expanding each point into a disk or ellipsoid with a certain radius, and blending the color and depth information of these disks or ellipsoids to form a smooth surface, resulting in a static background image. Another option is PBR voxelization, converting 3D point cloud data into voxel representation and combining it with physically based rendering (PBR) principles to calculate the lighting response of each voxel, thereby generating a static background image with realistic lighting and material properties. By employing these rendering methods for different types of static background representations, millions of point clouds can be efficiently rendered onto the screen, forming a static background image with realistic lighting and texture that covers the entire field of view.

[0052] Step 2032: Render the standard 3D model based on the second rendering pipeline to obtain a symbolic dynamic foreground.

[0053] In some examples, different standard 3D models can be used for different types of dynamic objects. For instance, if the dynamic object is a vehicle, the standard 3D model is a car model with uniform proportions and a simplified appearance, and its color can be set to a striking blue to distinguish it from the static background; if the dynamic object is a pedestrian, the standard 3D model can be a simplified human silhouette model, and the color can be orange to improve recognizability; for cyclists (such as bicycle riders and electric vehicle riders), the standard 3D model can be designed as a simplified combination model containing the cyclist and the corresponding vehicle, and the color can be green.

[0054] In some embodiments, during the rendering process, the second rendering pipeline automatically matches and calls the corresponding standard 3D model based on the category information of the dynamic object, and renders the standard 3D model of different dynamic objects at their corresponding spatial positions in the static background image through a symbolic SR renderer. It can also assign them a motion state consistent with the real dynamic object, thus obtaining a symbolic foreground. In this way, the visual style of the aforementioned standard 3D model remains consistent with the overall design language of the human-computer interaction interface, making it clear and easy to read.

[0055] In some embodiments, step 2032 may specifically include: processing the target image frame sequence based on a preset algorithm to obtain dynamic object information in the current surrounding environment; and rendering a standard 3D model based on the dynamic object information and a second rendering pipeline to obtain a symbolic dynamic foreground.

[0056] For example, the aforementioned preset algorithm may include at least one of the following: a bird's eye view (BEV) perception algorithm based on multi-camera fusion, and a target detection algorithm. For instance, the BEV perception algorithm may be a bird's eye view transformer (BEVFormer) or its variants; the target detection algorithm may be a Transformer-based DETR or other possible algorithms, etc.

[0057] In some examples, a multi-camera fusion-based bird's-eye view perception algorithm can be used to integrate target image frame sequences into a unified feature representation, which can then be converted into a BEV representation and the positional association between static background and dynamic objects can be extracted. Alternatively, object detection algorithms can be used to detect dynamic objects (such as vehicles and pedestrians) in the BEV representation to obtain dynamic object information for all dynamic objects in the vehicle's current surrounding environment. This dynamic object information can include the dynamic object's unique identifier (ID), position, speed, and orientation. For example, if the dynamic object is a vehicle, the dynamic object information could include the vehicle's category, ID, position, speed, and orientation.

[0058] In some examples, the placement coordinates of the standard 3D model are determined in the 3D coordinate system of the static background image based on the position information in the dynamic object information, ensuring that its relative position in space is consistent with that of the real dynamic object. Alternatively, the standard 3D model can be rotated based on the orientation information in the dynamic object information to match the direction of travel or movement of the real dynamic object. For the velocity information in the dynamic object information, it can be visually presented by overlaying velocity vector arrows on the standard 3D model or displaying velocity values ​​next to the model; the length and direction of the arrows correspond to the magnitude and direction of the velocity, respectively. If the dynamic object information also includes the dynamic object's ID, the ID can be rendered above or near the standard 3D model to distinguish and track different dynamic objects. During rendering, the second rendering pipeline calls the standard 3D model that matches the dynamic object category and, combined with the aforementioned position, orientation, velocity, and ID information, uses a symbolic SR renderer to draw the standard 3D model in a preset visual style (such as simplified geometry and vibrant colors) at the corresponding position in the static background image, ultimately forming a clear and easily recognizable symbolic dynamic foreground. Thus, the symbolic dynamic foreground obtained by combining dynamic object information with the rendering of a standard 3D model is more consistent with the real dynamic state of real dynamic objects in real physical space.

[0059] Step 2033: The static background image is fused with the symbolic dynamic foreground to obtain the target image.

[0060] In some examples, the static background images rendered separately can be merged with the symbolic dynamic foreground during the final rendering stage. Specifically, during rendering, the static background image and the symbolic dynamic foreground are placed in the same rendering space and share the same 3D coordinate system. A depth test mechanism is used to compare the depth values ​​of each pixel in the static background image and the symbolic dynamic foreground. When the depth value of a pixel in the symbolic dynamic foreground is less than the depth value of the corresponding pixel in the static background image, it indicates that the dynamic foreground pixel is in the foreground and should be displayed first, thus occluding the static background behind it. Conversely, if the depth value of the dynamic foreground pixel is greater than the depth value of the static background pixel, the dynamic foreground pixel is occluded by the static background and not displayed. The final generated target image shows a clearly defined symbolic vehicle driving in a road and building environment full of realistic details. Thus, this depth test-based fusion method achieves seamless fusion of the symbolic dynamic foreground and the static background image, ensuring the realism and consistency of the spatial relationship between the symbolic dynamic foreground and the static background image, making the dynamic object appear to exist naturally within the static environment. This creates a new in-vehicle augmented reality (AR) visual experience that is both realistic and clear.

[0061] Based on the above embodiments, since a static background image can be obtained by rendering a static background representation using the first rendering pipeline, and a symbolic dynamic foreground can be obtained by rendering a standard 3D model using the second rendering pipeline, and then the static background image and the symbolic dynamic foreground are merged to obtain the target image, the final target image retains the realistic details and environmental atmosphere of the static background, while highlighting key dynamic object information through a clear and eye-catching symbolic dynamic foreground, achieving an effective combination of realism and information readability. This provides drivers with rich real-world environmental context while ensuring rapid identification of key dynamic information such as surrounding vehicles and pedestrians, avoiding information overload. This differentiated display method helps drivers focus their attention on the most important dynamic traffic participants, improving driving safety. Moreover, compared to rendering static and dynamic objects using the same pipeline, this method uses differentiated rendering for the rendering target, which improves rendering efficiency and saves computing power.

[0062] In some embodiments, the target is dynamically represented as a standard three-dimensional model; such as Figure 5 As shown above, in the above Figure 2 Based on the illustrated embodiment, prior to step 203 above, the three-dimensional reconstruction method provided in this disclosure may further include the following steps: Step 204: The 3D scene representation is segmented based on a preset deep learning network to obtain the first semantic label and the second semantic label of the 3D scene representation.

[0063] The first semantic label is used to indicate the static background representation, and the second semantic label is used to indicate the dynamic object representation.

[0064] For example, the deep learning network described above can be used to extract local or global features from a 3D scene representation to output a semantic label assigned to each input point. Dynamic objects (e.g., vehicles, pedestrians) typically possess characteristic shape, size, and surface properties; static environments (e.g., roads, buildings, etc.) typically possess features of large or continuous surfaces. For instance, the deep learning network can be a PointNet network or its improved versions, such as PointNet++, Point-Transformer, or MinkowskiNet.

[0065] In some examples, a 3D scene representation can be input into a pre-defined deep learning network. The feature extraction layer of this network performs multilayer perceptron processing on the input 3D scene representation, capturing the local geometric features and global contextual information of the 3D point cloud. Sampling and grouping layers are used to downsample and group feature points into neighborhoods, progressively improving the feature abstraction level. Finally, a classification layer assigns a semantic label to each point cloud. If the deep learning network identifies a 3D point cloud as a static background representation in the 3D scene representation based on its features, a first semantic label representing the static background is output. If the deep learning network identifies a 3D point cloud as a dynamic object representation in the 3D scene representation based on its features, a second semantic label representing the dynamic object is output. In this way, accurate distinction between static and dynamic elements in a 3D scene can be achieved. Furthermore, based on the first and second semantic labels, static background representations can be filtered out from the 3D scene representation.

[0066] In some embodiments, step 204 above may specifically include: obtaining scene change information of historical time series; and using the scene change information to segment the three-dimensional scene representation based on a deep learning network to obtain a first semantic label and a second semantic label.

[0067] In some examples, the aforementioned scene change information refers to historical 3D point clouds or feature information of historical 3D point clouds; historical time series refers to the timestamps of historical 3D point clouds, for example, the historical time series is 0.5 seconds past. When performing semantic label prediction, deep learning networks not only rely on the features of the currently acquired 3D scene representation, but also combine the scene change information of historical time series. By comparing changes under consecutive timestamp scenes, deep learning networks can identify point clouds whose displacement exceeds a preset threshold or whose morphological changes conform to the motion laws of dynamic objects, that is, identify dynamic objects that have undergone displacement. Point clouds that have not undergone displacement or morphological changes are considered static backgrounds, and thus output the first semantic label and the second semantic label. In this way, the above method can more reliably identify objects that have undergone displacement, further improving the accuracy of static background and dynamic object segmentation, especially for the recognition of slow-moving dynamic objects or dynamic elements that are easily confused with static backgrounds (such as slowly moving vehicles or walking pedestrians).

[0068] Based on the above embodiments, since the three-dimensional scene representation can be segmented based on a preset deep learning network to obtain the first semantic label and the second semantic label of the three-dimensional scene representation, the separation of dynamic elements and static elements is achieved.

[0069] In some embodiments, such as Figure 6 As shown above, in the above Figure 2Based on the illustrated embodiment, step 201 above may specifically include the following steps: Step 2011: Synchronously acquire the current surrounding environment using an image sensor to obtain an initial image frame sequence.

[0070] In some examples, the aforementioned image sensors refer to multiple cameras deployed at different locations around the vehicle, such as the front, left, right, rear, left rear, and right rear of the vehicle. A surround-view camera system can be constructed using multiple cameras (e.g., fisheye cameras, wide-angle cameras, etc.) deployed around the vehicle. These cameras simultaneously capture images of the environment around the vehicle in different directions (e.g., forward, backward, left-front, right-front, left-rear, right-rear, etc.) at the same preset frame rate (e.g., 30fps), thereby acquiring an initial image frame sequence containing multiple viewpoints. Synchronous acquisition can be achieved through hardware trigger signals (e.g., transmitted synchronization pulses) or network time protocols, ensuring precise synchronization of exposure and frame capture at the microsecond (μs) level for all cameras. This guarantees temporal consistency of image frames from different viewpoints in the initial image frame sequence, thus providing temporally consistent multi-view image data for 3D reconstruction.

[0071] Step 2012: Obtain the intrinsic parameters of the image sensor.

[0072] In some examples, during the vehicle manufacturing phase, specialized calibration equipment such as checkerboard patterns is used to calibrate the intrinsic parameters of each image sensor to obtain these parameters. The intrinsic parameters of each image sensor may include focal length, principal point coordinates, and distortion coefficients (e.g., radial and tangential distortion coefficients). These parameters can be used to perform real-time distortion correction on the captured raw images. The calibrated intrinsic parameters can be stored in the vehicle's local storage. When 3D reconstruction is required during vehicle operation, the intrinsic parameters of each image sensor can be retrieved from the local storage.

[0073] Step 2013: Perform distortion correction on the initial image frame sequence based on intrinsic parameters to obtain the target image frame sequence.

[0074] In some examples, for each initial image frame in the initial image frame sequence, the distortion offset of the distorted point in each initial image frame can be calculated using the distortion coefficients in the intrinsic parameters. Then, the coordinates of the distortion-free pixel corresponding to that distorted point are obtained by using an anti-distortion model based on the distortion offset. By performing the above distortion correction process on all image frames in the initial image frame sequence, the original distorted initial image frames are corrected to normal images that conform to the perspective projection rules, thus obtaining a target image frame sequence that eliminates the influence of lens distortion. This provides a more accurate image data foundation for subsequent 3D reconstruction.

[0075] Based on the above embodiments, since the current surrounding environment can be synchronously acquired by the image sensor to obtain an initial image frame sequence, and the initial image frame sequence can be distorted based on the acquired intrinsic parameters of the image sensor to obtain the target image frame sequence, the temporal consistency of the target image frame sequences from different perspectives can be ensured, and the image distortion caused by the inherent optical characteristics of the camera lens can be eliminated, so that the pixels in each image frame can accurately correspond to the positions in the real physical space, thereby providing an accurate data foundation for the subsequent three-dimensional reconstruction process.

[0076] Exemplary device Figure 7 This is a schematic diagram of a three-dimensional reconstruction apparatus for a vehicle's real-time environment, provided as an exemplary embodiment of this disclosure. The apparatus can be installed on an object such as a vehicle to perform the three-dimensional reconstruction method for the vehicle's real-time environment according to any of the embodiments described above.

[0077] like Figure 7 As shown, the aforementioned device 300 may include: a first acquisition module 301, which can be used to acquire a sequence of target image frames of the current surrounding environment of the vehicle; wherein the target image frame sequence includes image frames from different shooting angles; a three-dimensional reconstruction module 302, which can be used to perform three-dimensional reconstruction based on the target image frame sequence to obtain a three-dimensional scene representation of the current surrounding environment; wherein the three-dimensional scene representation includes a static background representation and a dynamic object representation; and an image rendering module 303, which can be used to render the static background representation and the target dynamic representation to obtain a target image corresponding to the current surrounding environment; wherein the target dynamic representation is the dynamic object representation or a preset standard three-dimensional model.

[0078] In one possible implementation, the aforementioned 3D reconstruction module 302 can be specifically used to: encode the target image frame sequence based on a preset encoder network to obtain a target latent vector describing the current surrounding environment; obtain the historical latent vector of the previous moment; and perform forward propagation on the target latent vector and the historical latent vector based on a preset decoder network to obtain the 3D scene representation.

[0079] In one possible implementation, the target dynamic representation is the standard three-dimensional model; the image rendering module 303 described above can be specifically used to: render the static background representation based on the first rendering pipeline to obtain a static background image; render the standard three-dimensional model based on the second rendering pipeline to obtain the symbolic dynamic foreground; and fuse the static background image with the symbolic dynamic foreground to obtain the target image.

[0080] In one possible implementation, the target is dynamically represented as the standard 3D model; the above apparatus may further include: a semantic segmentation module, which can be used to segment the 3D scene representation based on a preset deep learning network to obtain a first semantic label and a second semantic label of the 3D scene representation; wherein, the first semantic label is used to indicate the static background representation, and the second semantic label is used to indicate the dynamic object representation.

[0081] In one possible implementation, the semantic segmentation module described above can be specifically used to: acquire scene change information from a historical time series; and, based on the deep learning network, segment the 3D scene representation using the scene change information to obtain the first semantic label and the second semantic label.

[0082] In one possible implementation, the image rendering module 303 described above can be specifically used to: process the target image frame sequence based on a preset algorithm to obtain dynamic object information in the current surrounding environment; and render based on the dynamic object information and the standard three-dimensional model to obtain the symbolized dynamic foreground.

[0083] In one possible implementation, the first acquisition module 301 described above can be specifically used to: synchronously acquire the current surrounding environment through an image sensor to obtain an initial image frame sequence; acquire the intrinsic parameters of the image sensor; and perform distortion correction processing on the initial image frame sequence based on the intrinsic parameters to obtain the target image frame sequence.

[0084] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.

[0085] Exemplary electronic devices Figure 8 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 111 and a memory 112.

[0086] The processor 111 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.

[0087] The memory 112 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 111 may execute one or more computer program instructions to implement the three-dimensional reconstruction method of the real-time vehicle environment and / or other desired functions of the various embodiments of this disclosure described above.

[0088] In one example, the electronic device 11 may also include an input device 113 and an output device 114, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0089] The input device 113 may include various sensors, including but not limited to: a distance sensor for detecting the distance between a target object and the vehicle; an image sensor for acquiring information about the vehicle's surrounding environment. In some examples, the input device may also include a pressure sensor for detecting seat pressure to determine the presence and location of passengers; a temperature sensor for monitoring the temperature inside the cabin; a humidity sensor for monitoring the humidity inside the cabin to assist in regulating the in-vehicle environment; an air quality sensor for monitoring in-vehicle air quality, such as carbon dioxide and volatile organic compounds (VOCs); a light sensor for detecting the intensity of light inside and outside the vehicle; an acceleration sensor for detecting changes in the vehicle's acceleration; a distance sensor for detecting the distance between the vehicle and other objects; a touchscreen sensor for interaction with the vehicle's infotainment system; biometric sensors, such as fingerprint recognition and facial recognition; a heart rate monitor for monitoring the driver's heart rate; a sound sensor for voice recognition and interaction to enable voice control; a seat sensor for monitoring seat usage, such as whether the seat is occupied and the passenger's body size; and wireless communication sensors, such as Bluetooth and Wi-Fi, for connecting to smart devices to achieve data transmission and remote control. In addition to the examples given above, the input device may include more or fewer sensors, which will not be elaborated here.

[0090] The output device 114 can output various information or signals to other hardware or devices, which may include displays, car audio systems, seats, windows, steering wheels, communication networks, and their connected remote output devices. The displays may include multiple different displays such as a driver's side display, a passenger side display, and a rear-seat display. The car audio system may include multiple speakers located in different positions within the vehicle cabin, and each display or speaker can operate independently.

[0091] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 11 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 11 may include any other suitable components depending on the specific application.

[0092] Exemplary computer program products and computer-readable storage media In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the three-dimensional reconstruction methods for the real-time vehicle environment of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0093] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0094] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the three-dimensional reconstruction methods for the real-time vehicle environment of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0095] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0096] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0097] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A method for three-dimensional reconstruction of a vehicle's real-time environment, comprising: Acquire a sequence of target image frames of the vehicle's current surrounding environment; wherein, the sequence of target image frames includes image frames from different shooting angles; Based on the target image frame sequence, a 3D reconstruction is performed to obtain a 3D scene representation of the current surrounding environment; wherein, the 3D scene representation includes a static background representation and a dynamic object representation; The static background representation and the dynamic target representation are rendered to obtain a target image corresponding to the current surrounding environment; wherein, the dynamic target representation is the dynamic object representation or a preset standard 3D model.

2. The method according to claim 1, wherein, The step of performing 3D reconstruction based on the target image frame sequence to obtain a 3D scene representation of the current surrounding environment includes: The target image frame sequence is encoded based on a preset encoder network to obtain a target potential vector that describes the current surrounding environment; Obtain the historical latent vector from the previous time step; The target latent vector and the historical latent vector are forward-propagated based on a preset decoder network to obtain the 3D scene representation.

3. The method according to claim 1, wherein, The target dynamic representation is the standard 3D model; the rendering of the static background representation and the target dynamic representation to obtain a target image corresponding to the current surrounding environment includes: The static background representation is rendered using the first rendering pipeline to obtain a static background image. The standard 3D model is rendered using the second rendering pipeline to obtain the symbolic dynamic foreground; The static background image is fused with the symbolic dynamic foreground to obtain the target image.

4. The method according to claim 1, wherein, The target dynamic representation is the standard 3D model; before rendering the static background representation and the target dynamic representation to obtain the target image corresponding to the current surrounding environment, the method further includes: The three-dimensional scene representation is segmented based on a preset deep learning network to obtain a first semantic label and a second semantic label for the three-dimensional scene representation; wherein, the first semantic label is used to indicate the static background representation, and the second semantic label is used to indicate the dynamic object representation.

5. The method according to claim 4, wherein, The three-dimensional scene representation is segmented using a preset deep learning network to obtain a first semantic label and a second semantic label for the three-dimensional scene, including: Obtain historical time series information on scene changes; Based on the deep learning network, the scene change information is used to segment the 3D scene representation to obtain the first semantic label and the second semantic label.

6. The method according to claim 3, wherein, The rendering of the standard 3D model based on the second rendering pipeline to obtain the symbolized dynamic foreground includes: The target image frame sequence is processed based on a preset algorithm to obtain dynamic object information in the current surrounding environment; The standard 3D model is rendered based on the dynamic object information and the second rendering pipeline to obtain the symbolic dynamic foreground.

7. The method according to claim 1, wherein, The acquisition of the target image frame sequence of the vehicle's current surrounding environment includes: The current surrounding environment is synchronously acquired using an image sensor to obtain an initial image frame sequence; Obtain the intrinsic parameters of the image sensor; The initial image frame sequence is subjected to distortion correction based on the intrinsic parameters to obtain the target image frame sequence.

8. A three-dimensional reconstruction device for a vehicle's real-time environment, comprising: The first acquisition module is used to acquire a sequence of target image frames of the vehicle's current surrounding environment; wherein, the sequence of target image frames includes image frames from different shooting angles; A 3D reconstruction module is used to perform 3D reconstruction based on the target image frame sequence to obtain a 3D scene representation of the current surrounding environment; wherein, the 3D scene representation includes a static background representation and a dynamic object representation; The image rendering module is used to render the static background representation and the target dynamic representation to obtain a target image corresponding to the current surrounding environment; wherein, the target dynamic representation is the dynamic object representation or a preset standard three-dimensional model.

9. A computer-readable storage medium storing a computer program for performing the three-dimensional reconstruction method for a real-time vehicle environment as described in any one of claims 1-7.

10. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the three-dimensional reconstruction method of the real-time vehicle environment as described in any one of claims 1-7.