Gaussian map generation method and device of vehicle scene, equipment, vehicle and medium

By acquiring multi-camera image data and using dense SLAM algorithm to optimize the 3D Gaussian ellipsoid, the problem of low accuracy of Gaussian maps in existing technologies is solved, and more accurate Gaussian maps of vehicle scenes are generated.

CN121458932APending Publication Date: 2026-02-03CHAFA FRIEDRICH SCHAFFEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511529792.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

In existing technologies, when generating Gaussian maps of vehicle scenes using the SfM method, the sparse point cloud results in a small number of 3D Gaussian ellipsoids, leading to low accuracy of the Gaussian map.

Method used

By acquiring Gaussian map generation data, including image pairs, intrinsic parameters, pose transformation matrices, and initial depth data from multiple cameras, the target depth data and camera pose data are determined using a multi-camera dense SLAM algorithm. The initial 3D Gaussian ellipsoid is then optimized to generate a Gaussian map of the vehicle scene.

Benefits of technology

The accuracy of Gaussian plots for vehicle scenes has been improved, generating dense 3D Gaussian ellipsoids and enhancing the matching degree of Gaussian plots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458932A_ABST
    Figure CN121458932A_ABST
Patent Text Reader

Abstract

The invention provides a Gaussian map generation method and device of a vehicle scene, equipment, a vehicle and a medium. In the method, after Gaussian map generation data are obtained, target depth data and target camera pose data corresponding to each reference image in the Gaussian map generation data are determined according to the Gaussian map generation data and a multi-camera dense SLAM algorithm; determining an initial 3D Gaussian ellipsoid according to each reference image and the target depth data and the target camera pose data corresponding to each reference image; and according to each reference image and the target depth data and the target camera pose data corresponding to each reference image, optimizing the initial 3D Gaussian ellipsoid to generate a Gaussian map of the vehicle scene. According to the scheme, the dense initial 3D Gaussian ellipsoid is determined through the Gaussian map generation data and the multi-camera dense SLAM algorithm, and the accuracy of the Gaussian map of the vehicle scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and in particular to a method, apparatus, device, vehicle, and medium for generating Gaussian graphs of vehicle scenes. Background Technology

[0002] During autonomous driving, vehicles need to generate high-precision 3D maps of the surrounding environment to perceive drivable areas, obstacle distribution, and terrain undulations, thereby generating safe paths. High-precision 3D maps can be generated by rendering a Gaussian map of the vehicle scene.

[0003] In existing technologies, Gaussian maps of vehicle scenes can be generated using 3D Gaussian Splatting (3DGS) technology. This involves first obtaining sparse point clouds from images captured by a camera using the Structure from Motion (SfM) method, then generating a sparse 3D Gaussian ellipsoid from the sparse point cloud, and finally optimizing the 3D Gaussian ellipsoid to obtain the Gaussian map of the vehicle scene.

[0004] In summary, existing methods for generating Gaussian maps of vehicle scenes obtain sparse point clouds through the SfM method, resulting in a limited number of 3D Gaussian ellipsoids and consequently, lower accuracy of the Gaussian maps for vehicle scenes. Summary of the Invention

[0005] The Gaussian map generation method, apparatus, device, vehicle, and medium for vehicle scenes provided in this application are intended to solve the problem in the prior art where the number of 3D Gaussian ellipsoids obtained by the SfM method is small, resulting in low accuracy of the Gaussian map for the vehicle scene.

[0006] In a first aspect, embodiments of this application provide a method for generating Gaussian graphs of vehicle scenes, including:

[0007] Gaussian graph generation data is obtained, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrix of each camera and the vehicle, and initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched.

[0008] Based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined;

[0009] Based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, determine the initial 3D Gaussian ellipsoid;

[0010] Based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, the initial 3D Gaussian ellipsoid is optimized to generate a Gaussian map of the vehicle scene.

[0011] In one possible implementation, determining the target depth data and target camera pose data corresponding to each reference image based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm includes:

[0012] Based on the Gaussian plot generation data and the multi-camera dense SLAM algorithm, the estimated depth data corresponding to each reference image is determined, as well as the reprojection loss value and depth loss value corresponding to each pixel in each reference image.

[0013] Based on the reprojection loss value and depth loss value corresponding to each pixel in each reference image, and the multi-camera dense SLAM algorithm, the target optical flow coordinates and weights corresponding to each pixel in each reference image are determined.

[0014] Based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image, a reprojection loss function and a depth loss function are generated.

[0015] Based on the reprojection loss function, the depth loss function, and the pose transformation matrix corresponding to each camera and the vehicle, a target loss function is generated.

[0016] Based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm, new initial depth data and new initial camera pose data are generated for each image in each image pair.

[0017] Update the first iteration count;

[0018] If the updated first iteration number is less than the first preset iteration threshold, repeat the above steps until the updated first iteration number is equal to the first preset iteration threshold. Then, use the new initial target depth data and the new initial camera pose data corresponding to each reference image as the target depth data and target camera pose data corresponding to each reference image.

[0019] In one possible implementation, determining the estimated depth data corresponding to each reference image, and the reprojection loss value and depth loss value corresponding to each pixel in each reference image, based on the Gaussian map generation data and the multi-camera dense SLAM algorithm, includes:

[0020] Based on each image pair, the initial depth data and initial camera pose data corresponding to each reference image, the initial camera pose data corresponding to each image to be matched, the intrinsic parameters of each camera, and the multi-camera dense SLAM algorithm, the reprojection loss value corresponding to each pixel in each reference image is determined.

[0021] Based on each reference image, the preset depth estimation model, and the initial depth data corresponding to each reference image, the estimated depth data corresponding to each reference image and the depth loss value corresponding to each pixel in each reference image are determined.

[0022] In one possible implementation, generating a reprojection loss function and a depth loss function based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image includes:

[0023] The reprojection loss function is generated based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, as well as the intrinsic parameters of each camera.

[0024] The depth loss function is generated based on the estimated depth data corresponding to each reference image.

[0025] In one possible implementation, generating new initial depth data and new initial camera pose data for each image in each image pair based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm includes:

[0026] The vehicle pose data and new initial depth data corresponding to each image in each image pair are generated based on the target loss function and the multi-camera dense SLAM algorithm.

[0027] For each reference image, based on the vehicle pose data and the pose transformation matrix between the camera that captured the reference image and the vehicle, new initial camera pose data corresponding to the reference image is obtained.

[0028] Based on the new initial camera pose data corresponding to each reference image, generate new initial camera pose data corresponding to each image to be matched.

[0029] In one possible implementation, optimizing the initial 3D Gaussian ellipsoid based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, to generate a Gaussian map of the vehicle scene, includes:

[0030] For each reference image, a rendering process is performed based on the target camera pose data corresponding to the reference image and the initial 3D Gaussian ellipsoid to obtain the rendered image corresponding to the reference image;

[0031] For each reference image, based on the reference image, the target depth data corresponding to the reference image, and the rendered image, generate the loss value, semantic loss value, and depth loss value corresponding to each pixel in the reference image;

[0032] The target loss value is calculated based on the photometric loss value, semantic loss value, and depth loss value corresponding to each pixel in each reference image.

[0033] Based on the target loss value and the 3D Gaussian splashing technique, the initial 3D Gaussian ellipsoid is updated to obtain a new 3D Gaussian ellipsoid.

[0034] Update the second iteration count;

[0035] If the updated second iteration number is less than the second preset iteration threshold, repeat the above steps until the updated second iteration number is equal to the second preset iteration threshold, and use the image composed of the new 3D Gaussian ellipsoid as the Gaussian map of the vehicle scene.

[0036] In one possible implementation, generating the photometric loss value, semantic loss value, and depth loss value for each pixel in the reference image based on the reference image, the target depth data corresponding to the reference image, and the rendered image includes:

[0037] Based on the color data of each pixel in the reference image and the rendered image, calculate the photometric loss value corresponding to each pixel in the reference image;

[0038] Based on the rendered image and the preset depth estimation model, generate depth data corresponding to the rendered image;

[0039] Based on the target depth data corresponding to the reference image and the depth data corresponding to the rendered image, calculate the depth loss value corresponding to each pixel in the reference image;

[0040] Based on the reference image, the rendered image, and the preset semantic model, the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image are obtained.

[0041] Based on the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image, the semantic loss value corresponding to each pixel in the reference image is calculated.

[0042] Secondly, embodiments of this application provide a Gaussian graph generation apparatus for a vehicle scene, comprising:

[0043] The acquisition module is used to acquire Gaussian map generation data, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrix of each camera and the vehicle, and initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched.

[0044] Processing module, used for:

[0045] Based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined;

[0046] Based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, determine the initial 3D Gaussian ellipsoid;

[0047] The generation module is used to optimize the initial 3D Gaussian ellipsoid based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, to generate a Gaussian map of the vehicle scene.

[0048] Thirdly, embodiments of this application provide an electronic device, including:

[0049] Processor, memory, communication interface;

[0050] The memory is used to store the executable instructions of the processor;

[0051] The processor is configured to execute the Gaussian graph generation method for the vehicle scene according to any one of the first aspects by executing the executable instructions.

[0052] Fourthly, embodiments of this application provide a vehicle, including a controller;

[0053] The controller is used to execute the Gaussian graph generation method for the vehicle scene as described in any of the first aspects above.

[0054] Fifthly, embodiments of this application provide a readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the Gaussian graph generation method for a vehicle scene as described in any of the first aspects.

[0055] In a sixth aspect, embodiments of this application provide a computer program product, including a computer program, which, when executed by a processor, is used to implement the Gaussian graph generation method for a vehicle scene as described in any of the first aspects.

[0056] The Gaussian map generation method, apparatus, device, vehicle, and medium for vehicle scenes provided in this application embodiment acquire Gaussian map generation data, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrices corresponding to each camera and the vehicle, and initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image and a matching image of the vehicle scene. Based on the Gaussian map generation data and a multi-camera dense SLAM algorithm, target depth data and target camera pose data corresponding to each reference image are determined. Then, based on each reference image and its corresponding target depth data and target camera pose data, an initial 3D Gaussian ellipsoid is determined. Finally, based on each reference image and its corresponding target depth data and target camera pose data, the initial 3D Gaussian ellipsoid is optimized to generate a Gaussian map of the vehicle scene. This solution, by using Gaussian map generation data and a multi-camera dense SLAM algorithm to determine a dense initial 3D Gaussian ellipsoid, can improve the accuracy of the Gaussian map of the vehicle scene. Attached Figure Description

[0057] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0058] Figure 1 A flowchart illustrating an embodiment of the Gaussian graph generation method for vehicle scenarios provided in this application;

[0059] Figure 2 A flowchart illustrating Embodiment 2 of the method for generating Gaussian graphs of vehicle scenes provided in this application;

[0060] Figure 3 A flowchart illustrating Embodiment 3 of the method for generating Gaussian graphs of vehicle scenes provided in this application;

[0061] Figure 4 A schematic diagram of the structure of an embodiment of the Gaussian graph generation device for vehicle scenes provided in this application;

[0062] Figure 5 This is a schematic diagram of the structure of an electronic device provided in this application.

[0063] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0064] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0065] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0066] With the continuous development of technology, vehicles are becoming capable of autonomous driving. During autonomous driving, vehicles need to generate high-precision 3D maps of the surrounding environment to perceive drivable areas, obstacle distribution, and terrain undulations, thereby generating safe paths. High-precision 3D maps can be generated by rendering a Gaussian map of the vehicle scene.

[0067] In existing technologies, Gaussian maps of vehicle scenes can be generated using 3D Gaussian Splatting (3DGS) technology. This involves first obtaining a sparse point cloud from images captured by a camera using the Structure from Motion (SfM) method. Then, a sparse 3D Gaussian ellipsoid is generated from this point cloud, and further optimization is performed to obtain the Gaussian map of the vehicle scene. However, because the SfM method generates a sparse point cloud, the number of 3D Gaussian ellipsoids is relatively small, leading to lower accuracy in the Gaussian map of the vehicle scene—that is, a lower degree of matching between the Gaussian map and the vehicle scene.

[0068] To address the problems existing in the prior art, the inventors, during their research on Gaussian map generation methods for vehicle scenes, discovered that to improve the accuracy of Gaussian maps for vehicle scenes, Gaussian map generation data can be obtained first. This data includes multiple image pairs captured by multiple cameras in the vehicle, the intrinsic parameters of each camera, the pose transformation matrix corresponding to each camera and the vehicle, and the initial depth data and initial camera pose data for each image in each image pair. Each image pair includes a reference image and an image to be matched for the vehicle scene. Then, based on the Gaussian map generation data and a multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined, thus establishing an initial 3D Gaussian ellipsoid. Next, based on each reference image and its corresponding target depth data and target camera pose data, the initial 3D Gaussian ellipsoid is optimized to generate the Gaussian map of the vehicle scene. By using the Gaussian map generation data and the multi-camera dense SLAM algorithm to determine a dense initial 3D Gaussian ellipsoid, the accuracy of the Gaussian map for the vehicle scene can be improved. Based on the above-mentioned inventive concept, a Gaussian graph generation scheme for the vehicle scene in this application was designed.

[0069] The execution entity of the Gaussian graph generation method for vehicle scenes in this application can be a controller in the vehicle, or a computer, server, vehicle terminal, etc. This application does not limit it. The controller is used as an example for explanation below.

[0070] The following provides an example illustrating the application scenarios of the Gaussian graph generation method for vehicle scenarios provided in this application.

[0071] For example, in this application scenario, an autonomous vehicle is driving in the wilderness, and a high-precision map needs to be built to plan its route. The vehicle is equipped with multiple cameras, each capturing images of the scene from different angles.

[0072] The controller in the vehicle first acquires Gaussian map generation data, which includes multiple image pairs captured by multiple cameras in the vehicle, the intrinsic parameters of each camera, the pose transformation matrix of each camera and the vehicle, and the initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched.

[0073] Then, based on the Gaussian image generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined; and based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, the initial 3D Gaussian ellipsoid is determined.

[0074] It should be noted that multi-camera dense SLAM can be algorithms such as DROID-SLAM (Differentiable Recurrent Optimization-Inspired Design Simultaneous Localization And Mapping) or other algorithms that can achieve similar functions.

[0075] Then, based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, the initial 3D Gaussian ellipsoid is optimized to generate a Gaussian map of the vehicle scene.

[0076] The controller renders the Gaussian graph of the vehicle scene to generate a high-precision 3D map, and then performs path planning based on the high-precision 3D map.

[0077] It should be noted that the above scenario is only an example of an application scenario provided by the embodiments of this application. The embodiments of this application do not limit the actual form of the various devices included in the scenario, nor do they limit the interaction method between devices. In the specific application of the solution, it can be set according to actual needs.

[0078] The technical solution of this application will now be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0079] Figure 1 This is a flowchart illustrating an embodiment of the Gaussian map generation method for vehicle scenes provided in this application. This embodiment describes how the controller generates a dense initial 3D Gaussian ellipsoid based on the Gaussian map generation data and a multi-camera dense SLAM algorithm, and then optimizes it to obtain the Gaussian map of the vehicle scene. The method in this embodiment can be implemented through software, hardware, or a combination of both. Figure 1 As shown, the method for generating the Gaussian graph of this vehicle scene specifically includes the following steps:

[0080] S101: Obtain Gaussian plot generation data.

[0081] In this step, in order to generate a Gaussian plot of the vehicle scene, the controller first acquires the Gaussian plot generation data.

[0082] The Gaussian graph generation data includes multiple image pairs captured by multiple cameras in the vehicle, the intrinsic parameters of each camera, the pose transformation matrix of each camera and the vehicle, and the initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched.

[0083] It should be noted that each camera in the vehicle captures multiple images within the same time window. For each camera, the images are sorted in chronological order of capture time, resulting in an image sequence. Each pair of adjacent images in the sequence is then considered an image pair; the earlier image in each pair serves as the reference image, and the later image serves as the image to be matched. Therefore, each image in the sequence, except for the first and last images, exists in two image pairs: one as the reference image and the other as the image to be matched. Specifically, the first image in the image sequence is the reference image, and the last image is the image to be matched.

[0084] For example, the image sequence is A, B, C, D, and the image pairs are (A, B), (B, C), and (C, D). Image A is the reference image in the first image pair; image B is the reference image in the second image pair and the image to be matched in the first image pair; image C is the reference image in the third image pair and the image to be matched in the second image pair; and image D is the image to be matched in the third image pair.

[0085] S102: Based on the Gaussian map generation data and the multi-camera dense SLAM algorithm, determine the target depth data and target camera pose data corresponding to each reference image.

[0086] In this step, after the controller obtains the Gaussian map generation data, in order to improve the number and accuracy of the subsequent initial 3D Gaussian ellipsoids, the target depth data and target camera pose data corresponding to each reference image are determined based on the Gaussian map generation data and the multi-camera dense SLAM algorithm.

[0087] The controller can determine the estimated depth data for each reference image, as well as the reprojection loss value and depth loss value for each pixel in each reference image, based on Gaussian map generation data and a multi-camera dense SLAM algorithm. Based on the reprojection loss value and depth loss value, the controller determines the target optical flow coordinates and weights for each pixel in each reference image. Then, combining the intrinsic parameters of each camera and the estimated depth data for each reference image, it generates a reprojection loss function and a depth loss function. Based on the reprojection loss function and depth loss function, the controller updates the initial depth data and initial camera pose data for each image. This process is repeated until a threshold number of iterations is reached, resulting in the target depth data and target camera pose data for each reference image.

[0088] Processing each pixel of each reference image increases the number of subsequent initial 3D Gaussian ellipsoids; the accuracy of subsequent initial 3D Gaussian ellipsoids can be improved through reprojection loss function and depth loss function.

[0089] S103: Determine the initial 3D Gaussian ellipsoid based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image.

[0090] In this step, after the controller obtains the target depth data and target camera pose data corresponding to each reference image, it combines each reference image to determine the initial 3D Gaussian ellipsoid.

[0091] The target depth data contains the depth value corresponding to each pixel in the reference image. The reference image is input into a preset semantic model to obtain the semantic data for each pixel in the reference image. The reference image is then input into a preset transparency detection model to obtain the transparency of each pixel in the reference image.

[0092] The preset semantic model is a pre-trained neural network model used to determine the semantic data of each pixel in an image. The preset transparency detection model is a pre-trained neural network model used to determine the transparency of each pixel in an image.

[0093] Specifically, for each pixel in each reference image, the initial position of the initial 3D Gaussian ellipsoid is determined based on the depth value corresponding to that pixel and the target camera pose data corresponding to that reference image. The color data of that pixel is used as the initial color data of the initial 3D Gaussian ellipsoid. The semantic data of that pixel is used as the initial semantic data of the initial 3D Gaussian ellipsoid. The transparency of that pixel is used as the transparency of the initial 3D Gaussian ellipsoid, and the initial shape data of the initial 3D Gaussian ellipsoid is set to default data. Based on the initial position, initial color data, initial semantic data, transparency, and initial shape data of the initial 3D Gaussian ellipsoid, an initial 3D Gaussian ellipsoid is generated.

[0094] For each pixel in each reference image, a corresponding initial 3D Gaussian ellipsoid can be generated. The large number of initial 3D Gaussian ellipsoids can improve the accuracy of the Gaussian map of the vehicle scene generated subsequently.

[0095] S104: Based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, optimize the initial 3D Gaussian ellipsoid to generate a Gaussian map of the vehicle scene.

[0096] In this step, after the controller generates an initial 3D Gaussian ellipsoid, it optimizes the initial 3D Gaussian ellipsoid based on each reference image and the target depth data and target camera pose data corresponding to each reference image to generate a Gaussian map of the vehicle scene.

[0097] The controller renders an initial 3D Gaussian ellipsoid based on the target pose data corresponding to the reference image, obtaining a rendered image corresponding to the reference image. Then, based on the reference image and its corresponding rendered image, it generates loss values, semantic loss values, and depth loss values ​​for each pixel in the reference image. Combining this with 3D Gaussian splashing technology, the initial 3D Gaussian ellipsoid is updated to obtain a new 3D Gaussian ellipsoid. This process is repeated until a threshold is reached, and the image formed by the final 3D Gaussian ellipsoid is used as the Gaussian map of the vehicle scene.

[0098] The Gaussian graph of the vehicle scene can then be used for suspension dynamics adaptation, path planning, scene image rendering, etc.

[0099] The Gaussian map generation method for vehicle scenes provided in this embodiment obtains Gaussian map generation data, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrices corresponding to each camera and the vehicle, and initial depth data and initial camera pose data for each image in each image pair. Each image pair includes a reference image and an image to be matched for the vehicle scene. Based on the Gaussian map generation data and a multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined. Then, based on each reference image and its corresponding target depth data and target camera pose data, an initial 3D Gaussian ellipsoid is determined. Finally, based on each reference image and its corresponding target depth data and target camera pose data, the initial 3D Gaussian ellipsoid is optimized to generate the Gaussian map of the vehicle scene. This solution improves the accuracy of the Gaussian map of the vehicle scene by determining a dense initial 3D Gaussian ellipsoid through Gaussian map generation data and a multi-camera dense SLAM algorithm.

[0100] Figure 2 This is a flowchart illustrating a second embodiment of the Gaussian map generation method for vehicle scenes provided in this application. Based on the above embodiments, this application describes how the controller determines the target depth data and target camera pose data corresponding to each reference image according to the Gaussian map generation data and the multi-camera dense SLAM algorithm. For example... Figure 2 As shown, the method for generating the Gaussian graph of this vehicle scene specifically includes the following steps:

[0101] S201: Based on the Gaussian map generation data and the multi-camera dense SLAM algorithm, determine the estimated depth data corresponding to each reference image, as well as the reprojection loss value and depth loss value corresponding to each pixel in each reference image.

[0102] In this step, after the controller obtains the Gaussian map generation data, it determines the estimated depth data corresponding to each reference image, as well as the reprojection loss value and depth loss value corresponding to each pixel in each reference image, based on the Gaussian map generation data and the multi-camera dense SLAM algorithm.

[0103] Specifically, based on the initial depth data and initial camera pose data corresponding to each image pair, each reference image, the initial camera pose data corresponding to each image to be matched, the intrinsic parameters of each camera, and the multi-camera dense SLAM algorithm, the reprojection loss value corresponding to each pixel in each reference image is determined.

[0104] The initial camera pose data includes an initial pose matrix. The initial depth data includes the depth value corresponding to each pixel, so the depth values ​​can be arranged according to the pixel arrangement to form an initial depth matrix.

[0105] The multi-camera dense SLAM algorithm is the DROID-SLAM algorithm. For each image pair, the DROID-SLAM algorithm includes a pixel matching method. This method processes the two images in the pair to determine the optical flow matching position matrix of each pixel in the reference image within the image to be matched. Based on the initial depth and initial pose matrices of the reference image, the initial pose matrix of the image to be matched, the intrinsic parameters of the camera that captured the image pair, and the original position matrix of each pixel in the reference image, the mapping position matrix of each pixel in the reference image within the image to be matched is determined.

[0106] Through formula Determine the mapping position matrix of a pixel in the reference image in the image to be matched, where, This represents the mapping position matrix of a pixel in the reference image to the image to be matched, where K represents the intrinsic parameters of the camera that captured the image pair. This represents the initial pose matrix corresponding to the reference image. Let D represent the initial pose matrix corresponding to the image to be matched, and let D represent the initial depth matrix corresponding to the reference image. This represents the original position matrix of that pixel in the reference image.

[0107] Furthermore, for each pixel in the reference image, the square of the norm of the difference between the optical flow matching position matrix and the mapping position matrix in the image to be matched is used as the reprojection loss value corresponding to that pixel.

[0108] The controller determines the estimated depth data and depth loss value for each reference image based on each reference image, the preset depth estimation model, and the initial depth data corresponding to each reference image.

[0109] For each reference image, the reference image is input into a preset depth estimation model to obtain the estimated depth data corresponding to the reference image. The estimated depth data includes the depth value corresponding to each pixel. For each pixel, the square of the difference between the initial depth data and the depth value of that pixel in the estimated depth data is used as the depth loss value corresponding to that pixel.

[0110] The preset depth estimation model is a pre-trained neural network model used to determine the depth value corresponding to each pixel in the image.

[0111] S202: Based on the reprojection loss value and depth loss value corresponding to each pixel in each reference image, and the multi-camera dense SLAM algorithm, determine the target optical flow coordinates and weights corresponding to each pixel in each reference image.

[0112] In this step, after the controller obtains the reprojection loss value and depth loss value corresponding to each pixel in each reference image, it determines the target optical flow coordinates and weights corresponding to each pixel in each reference image based on the reprojection loss value and depth loss value corresponding to each pixel in each reference image and the multi-camera dense SLAM algorithm.

[0113] Specifically, for each pixel in each reference image, the sum of the reprojection loss value and the depth loss value corresponding to that pixel is used as the residual value of that pixel.

[0114] The multi-camera dense SLAM algorithm is the DROID-SLAM algorithm. The DROID-SLAM algorithm has a method to determine the target optical flow coordinates and weights corresponding to each pixel in the reference image based on the residual values. Therefore, based on the reference image, the image to be matched in the same image pair as the reference image, and the residual values ​​of each pixel in the reference image, this method is used to process and obtain the target optical flow coordinates and weights corresponding to each pixel in the reference image.

[0115] S203: Generate a reprojection loss function and a depth loss function based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image.

[0116] In this step, after the controller obtains the target optical flow coordinates and weights corresponding to each pixel in each reference image, it combines the intrinsic parameters of each camera and the estimated depth data corresponding to each reference image to generate a reprojection loss function and a depth loss function.

[0117] Specifically, a reprojection loss function is generated based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, as well as the intrinsic parameters of each camera.

[0118] First, the undetermined depth matrix and undetermined camera pose matrix for each reference image, and the undetermined camera pose matrix for each image to be matched are set as unknowns. Then, for each image pair, based on the undetermined depth matrix and undetermined pose matrix for the reference image, the undetermined pose matrix for the image to be matched, the intrinsic parameters of the camera that captured the image pair, and the original position matrix of each pixel in the reference image, the undetermined mapping position matrix of each pixel in the reference image in the image to be matched is determined.

[0119] For each pixel in the reference image, an optical flow coordinate matrix is ​​constructed from the target optical flow coordinates corresponding to that pixel. The square of the norm of the difference between the optical flow coordinate matrix corresponding to that pixel and the matrix of the mapping position to be determined is used as the first projection loss function for that pixel. The product of the first projection loss function and the weights is used as the second projection loss function for that pixel.

[0120] The second projection loss function corresponding to each pixel in the reference image is summed to obtain the third projection loss function corresponding to the reference image.

[0121] The reprojection loss function is obtained by summing the third reprojection loss functions corresponding to each reference image.

[0122] A depth loss function is generated based on the estimated depth data corresponding to each reference image. That is, for each reference image, the depth values ​​in the estimated depth data corresponding to the reference image are arranged according to pixels to obtain an estimated depth matrix. Then, the square of the norm of the difference between the estimated depth matrix and the initial depth matrix corresponding to the reference image is used as the depth loss sub-function for that reference image.

[0123] Then, the depth loss sub-functions corresponding to each reference image are summed to obtain the depth loss function.

[0124] S204: Generate the target loss function based on the reprojection loss function, the depth loss function, and the pose transformation matrix corresponding to each camera and the vehicle.

[0125] In this step, after obtaining the reprojection loss function and the depth loss function, the controller combines the pose transformation matrix corresponding to each camera and the vehicle to generate the target loss function.

[0126] Specifically, the camera pose matrix to be determined in the reprojection loss function is an unknown quantity. In order to reduce the number of unknowns and improve computational efficiency, different camera pose matrices to be determined can be converted into the same vehicle pose matrix to be determined.

[0127] The vehicle pose matrix to be determined is set as an unknown. For each camera pose matrix to be determined in the reprojection loss function, this matrix is ​​replaced with the product of the vehicle pose matrix to be determined and the target pose transformation matrix. The target pose transformation matrix is ​​the pose transformation matrix corresponding to the target camera and the vehicle, and the target camera is the camera that captured the image corresponding to this pose matrix. After each pose matrix is ​​replaced, the reprojection update loss function is obtained. The reprojection update loss function is a function of the vehicle pose matrix and the depth matrix to be determined.

[0128] The target loss function is the sum of the product of the reprojection update loss function and the first preset weight, and the product of the depth loss function and the second preset weight.

[0129] It should be noted that the first preset weight and the second preset weight can be 0.1, 0.2, 0.3, 0.4, 0.6, 0.8, 0.9, etc. The embodiments of this application do not limit the first preset weight and the second preset weight, and can be determined according to the actual situation.

[0130] S205: Based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm, generate new initial depth data and new initial camera pose data for each image in each image pair.

[0131] In this step, after obtaining the target loss function, the controller generates new initial depth data and new initial camera pose data for each image in each image pair based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm. New Gaussian map generation data is also obtained.

[0132] Specifically, based on the target loss function and the multi-camera dense SLAM algorithm, vehicle pose data and new initial depth data corresponding to each image in each image pair are generated.

[0133] The multi-camera dense SLAM algorithm is the DROID-SLAM algorithm. Since the DROID-SLAM algorithm has a method to determine the update value corresponding to the initial value based on the loss function and the initial value, the initial camera pose data corresponding to each reference image and the initial camera pose data corresponding to each image to be matched are combined with the pose transformation matrix corresponding to each camera and the vehicle to determine the initial vehicle pose data. Then, based on the initial vehicle pose data, the initial depth data corresponding to each reference image, and the target loss function, this method is used to process the data to obtain the vehicle pose data and the new initial depth data corresponding to each reference image.

[0134] Then, for each reference image, based on the vehicle pose data and the pose transformation matrix between the camera that captured the reference image and the vehicle, a new initial camera pose data corresponding to the reference image is obtained.

[0135] For each image to be matched, new initial camera pose data corresponding to the image to be matched is obtained based on the vehicle pose data and the pose transformation matrix of the camera that captured the image and the vehicle.

[0136] Based on the new initial camera pose data corresponding to each reference image, generate new initial camera pose data corresponding to each image to be matched.

[0137] Since each image in the image sequence, except for the first and last images, is both a reference image and an image to be matched, the new initial depth data of each image to be matched, except for the last image, is also known when the new initial depth data corresponding to each reference image is known. Therefore, the initial depth data of the last image to be matched in the image sequence is used as its new initial depth data.

[0138] S206: Update the first iteration count.

[0139] In this step, after the controller obtains the new initial camera pose data and the new initial depth data, in order to determine whether the initial camera pose data and the initial depth data still need to be updated, the first iteration number needs to be updated, that is, the first iteration number is incremented by one.

[0140] S207: Determine whether the number of the first iterations after the update is less than the first preset iteration threshold; if the number of the first iterations after the update is less than the first preset iteration threshold, then execute step S201; if the number of the first iterations after the update is equal to the first preset iteration threshold, then execute step S208.

[0141] In this step, after the controller updates the first iteration count, in order to determine whether the initial camera pose data and initial depth data still need to be updated, it is determined whether the updated first iteration count is less than the first preset iteration threshold.

[0142] It should be noted that the first preset iteration threshold can be 10, 100, 1000, etc. This application embodiment does not limit the first preset iteration threshold, and it can be determined according to the actual situation.

[0143] If the number of the first iteration after the update is less than the first preset iteration threshold, it means that further updates are needed, and the process returns to step S201. That is, the above steps are repeated until the number of the first iteration after the update equals the first preset iteration threshold.

[0144] S208: Use the new initial target depth data and new initial camera pose data corresponding to each reference image as the target depth data and target camera pose data corresponding to each reference image.

[0145] In this step, if the controller determines that the number of the first iteration after the update is equal to the first preset iteration threshold, it means that no further updates are needed. Then, the new initial target depth data and the new initial camera pose data corresponding to each reference image are used as the target depth data and target camera pose data corresponding to each reference image.

[0146] The Gaussian map generation method for vehicle scenes provided in this embodiment transforms multiple unknowns into a single unknown by converting the pose transformation matrix corresponding to the camera and the vehicle. This reduces computational load, improves computational efficiency, and increases the efficiency of obtaining target depth data and target camera pose data. Furthermore, it enhances the geometric consistency of the target camera pose data. Generating target optical flow coordinates and weights using reprojection loss and depth loss values ​​improves their accuracy. Using a target loss function generated from the reprojection loss function and depth loss function, new initial depth data and new initial camera pose data are generated, improving their accuracy and mitigating the depth estimation drift problem in textureless regions using traditional SLAM.

[0147] Figure 3 This is a flowchart illustrating a third embodiment of the Gaussian map generation method for vehicle scenes provided in this application. Based on the above embodiments, this application describes how the controller optimizes the initial 3D Gaussian ellipsoid based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, to generate a Gaussian map of the vehicle scene. For example... Figure 3 As shown, the method for generating the Gaussian graph of this vehicle scene specifically includes the following steps:

[0148] S301: For each reference image, perform rendering processing based on the target camera pose data corresponding to the reference image and the initial 3D Gaussian ellipsoid to obtain the rendered image corresponding to the reference image.

[0149] In this step, after the controller obtains the initial 3D Gaussian ellipsoid, it needs to perform rendering processing on each reference image based on the target camera pose data corresponding to the reference image and the initial 3D Gaussian ellipsoid to obtain the rendered image corresponding to the reference image. This rendered image is the image from the viewpoint corresponding to the target camera pose data of the reference image.

[0150] It should be noted that the 3D Gaussian splashing technology includes a rendering method, which can be used to render the initial 3D Gaussian ellipsoid to obtain a rendered image.

[0151] S302: For each reference image, based on the reference image, the target depth data corresponding to the reference image, and the rendered image, generate the loss value, semantic loss value, and depth loss value corresponding to each pixel in the reference image.

[0152] In this step, after the controller obtains the rendered image corresponding to each reference image, for each reference image, based on the reference image, the target depth data corresponding to the reference image, and the rendered image, it generates the loss value, semantic loss value, and depth loss value corresponding to each pixel in the reference image.

[0153] Specifically, based on the color data of each pixel in the reference image and the rendered image, the photometric loss value corresponding to each pixel in the reference image is calculated. The color data can be constructed into an RGB matrix. For each pixel, the square of the norm of the difference between the RGB matrix of that pixel in the reference image and the rendered image is used as the photometric loss value corresponding to that pixel.

[0154] Based on the rendered image and a preset depth estimation model, depth data corresponding to the rendered image is generated. In other words, the rendered image is input into the preset depth estimation model to obtain the depth data corresponding to the rendered image.

[0155] Then, based on the target depth data corresponding to the reference image and the depth data corresponding to the rendered image, the depth loss value corresponding to each pixel in the reference image is calculated. Both the target depth data and the depth data corresponding to the rendered image contain the depth value corresponding to each pixel. Therefore, for each pixel, the square of the difference between the depth loss value corresponding to that pixel in the target depth data and the depth data corresponding to the rendered image is taken as the depth loss value corresponding to that pixel.

[0156] Based on the reference image, the rendered image, and the preset semantic model, the semantic data corresponding to the reference image and the rendered image are obtained. The reference image is input into the preset semantic model to obtain the semantic data corresponding to the reference image. The rendered image is input into the preset semantic model to obtain the semantic data corresponding to the rendered image.

[0157] Based on the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image, the semantic loss value corresponding to each pixel in the reference image is calculated. The semantic data includes the semantic matrix corresponding to each pixel. Therefore, for each pixel, the square of the difference between the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image is used as the semantic loss value corresponding to that pixel.

[0158] S303: Calculate the target loss value based on the photometric loss value, semantic loss value, and depth loss value corresponding to each pixel in each reference image.

[0159] In this step, after the controller obtains the photometric loss value, semantic loss value, and depth loss value corresponding to each pixel in each reference image, it can calculate the target loss value.

[0160] Specifically, the photometric loss value corresponding to each pixel in each reference image is summed to obtain the first loss value; the semantic loss value corresponding to each pixel in each reference image is summed to obtain the second loss value; and the depth loss value corresponding to each pixel in each reference image is summed to obtain the third loss value.

[0161] The sum of the products of the first loss value and the third preset weight, the second loss value and the fourth preset weight, and the third loss value and the fifth preset weight is taken as the target loss value.

[0162] It should be noted that the third, fourth, and fifth preset weights can be 0.1, 0.2, 0.3, 0.4, 0.6, 0.8, 0.9, etc. The embodiments of this application do not limit the third, fourth, and fifth preset weights, and they can be determined according to the actual situation.

[0163] S304: Based on the target loss value and the 3D Gaussian splashing technique, update the initial 3D Gaussian ellipsoid to obtain a new 3D Gaussian ellipsoid.

[0164] In this step, after the controller obtains the target loss value, it updates the initial 3D Gaussian ellipsoid based on the target loss value and the 3D Gaussian splashing technique to obtain a new 3D Gaussian ellipsoid.

[0165] Since the 3D Gaussian splashing technique has a method to update the 3D Gaussian ellipsoid based on the loss value, the initial 3D Gaussian ellipsoid can be updated according to this method to obtain a new 3D Gaussian ellipsoid.

[0166] It should be noted that updating a 3D Gaussian ellipsoid can be done by updating only its position, color data, semantic data, transparency, and shape data. Alternatively, it can update all of these data, including the number of 3D Gaussian ellipsoids.

[0167] S305: Update the second iteration count.

[0168] In this step, after the controller obtains the new 3D Gaussian ellipsoid, in order to determine whether to continue updating and optimizing the new 3D Gaussian ellipsoid, it is necessary to update the second iteration number, that is, to increment the second iteration number by one.

[0169] S306: Determine whether the updated second iteration number is less than the second preset iteration threshold; if the updated second iteration number is less than the second preset iteration threshold, then execute step S301; if the updated second iteration number is equal to the second preset iteration threshold, then execute step S307.

[0170] In this step, after the controller updates the second iteration number, in order to determine whether to continue updating and optimizing the new 3D Gaussian ellipsoid, it is determined whether the updated second iteration number is less than the second preset iteration threshold.

[0171] It should be noted that the second preset iteration threshold can be 100, 1000, 10000, etc. This application embodiment does not limit the second preset iteration threshold, and it can be determined according to the actual situation.

[0172] If the number of iterations after the update is less than the second preset iteration threshold, it means that further updates and optimizations are needed, and the process returns to step S301. That is, the above steps are repeated until the number of iterations after the update equals the second preset iteration threshold.

[0173] S307: Use the image formed by the new 3D Gaussian ellipsoid as the Gaussian plot of the vehicle scene.

[0174] In this step, if the controller determines that the updated second iteration number is equal to the second preset iteration threshold, it means that no further updates and optimizations are needed, and the image composed of the new 3D Gaussian ellipsoid is used as the Gaussian map of the vehicle scene.

[0175] The Gaussian map generation method for vehicle scenes provided in this embodiment updates the 3D Gaussian ellipsoid by calculating photometric loss, semantic loss, and depth loss values, thereby improving the accuracy of the 3D Gaussian ellipsoid and Gaussian map. The depth loss value helps avoid erroneous dilation of the Gaussian ellipsoid in occluded areas. The 3D Gaussian ellipsoid contains semantic data, which enhances the ability to understand the scene when using the Gaussian map subsequently.

[0176] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0177] Figure 4 This is a schematic diagram of the structure of an embodiment of the Gaussian diagram generation device for vehicle scenes provided in this application. Figure 4 As shown, the Gaussian graph generation device 40 for the vehicle scene includes:

[0178] The acquisition module 41 is used to acquire Gaussian map generation data. The Gaussian map generation data includes multiple image pairs captured by multiple cameras in the vehicle, the intrinsic parameters of each camera, the pose transformation matrix of each camera and the vehicle, and the initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched.

[0179] Processing module 42 is used for:

[0180] Based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined;

[0181] Based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, determine the initial 3D Gaussian ellipsoid;

[0182] The generation module 43 is used to optimize the initial 3D Gaussian ellipsoid based on each reference image and the target depth data and target camera pose data corresponding to each reference image, and generate a Gaussian map of the vehicle scene.

[0183] Furthermore, processing module 42 is specifically used for:

[0184] Based on the Gaussian plot generation data and the multi-camera dense SLAM algorithm, the estimated depth data corresponding to each reference image is determined, as well as the reprojection loss value and depth loss value corresponding to each pixel in each reference image;

[0185] Based on the reprojection loss value and depth loss value corresponding to each pixel in each reference image, and the multi-camera dense SLAM algorithm, the target optical flow coordinates and weights corresponding to each pixel in each reference image are determined.

[0186] Based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image, a reprojection loss function and a depth loss function are generated.

[0187] The target loss function is generated based on the reprojection loss function, the depth loss function, and the pose transformation matrix corresponding to each camera and the vehicle.

[0188] Based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm, new initial depth data and new initial camera pose data are generated for each image in each image pair.

[0189] Update the first iteration count;

[0190] If the updated first iteration number is less than the first preset iteration threshold, repeat the above steps until the updated first iteration number is equal to the first preset iteration threshold. Then, use the new initial target depth data and the new initial camera pose data corresponding to each reference image as the target depth data and target camera pose data corresponding to each reference image.

[0191] Furthermore, processing module 42 is specifically used for:

[0192] Based on each image pair, the initial depth data and initial camera pose data corresponding to each reference image, the initial camera pose data corresponding to each image to be matched, the intrinsic parameters of each camera, and the multi-camera dense SLAM algorithm, determine the reprojection loss value of each pixel in each reference image.

[0193] Based on each reference image, the preset depth estimation model, and the initial depth data corresponding to each reference image, the estimated depth data corresponding to each reference image and the depth loss value corresponding to each pixel in each reference image are determined.

[0194] Furthermore, processing module 42 is specifically used for:

[0195] Based on the target loss function and the multi-camera dense SLAM algorithm, vehicle pose data and new initial depth data corresponding to each image in each image pair are generated;

[0196] For each reference image, based on the vehicle pose data and the pose transformation matrix between the camera that captured the reference image and the vehicle, new initial camera pose data corresponding to the reference image is obtained.

[0197] Based on the new initial camera pose data corresponding to each reference image, generate new initial camera pose data corresponding to each image to be matched.

[0198] Furthermore, module 43 is specifically used for:

[0199] For each reference image, rendering is performed based on the target camera pose data and the initial 3D Gaussian ellipsoid corresponding to the reference image to obtain the rendered image corresponding to the reference image.

[0200] For each reference image, based on the reference image, the target depth data corresponding to the reference image, and the rendered image, generate the loss value, semantic loss value, and depth loss value for each pixel in the reference image;

[0201] The target loss value is calculated based on the photometric loss value, semantic loss value, and depth loss value corresponding to each pixel in each reference image.

[0202] Based on the target loss value and the 3D Gaussian splashing technique, the initial 3D Gaussian ellipsoid is updated to obtain a new 3D Gaussian ellipsoid.

[0203] Update the second iteration count;

[0204] If the updated second iteration number is less than the second preset iteration threshold, repeat the above steps until the updated second iteration number is equal to the second preset iteration threshold, and use the image composed of the new 3D Gaussian ellipsoid as the Gaussian map of the vehicle scene.

[0205] Furthermore, module 43 is specifically used for:

[0206] Based on the color data of each pixel in the reference image and the rendered image, calculate the photometric loss value corresponding to each pixel in the reference image;

[0207] Based on the rendered image and the preset depth estimation model, generate the depth data corresponding to the rendered image;

[0208] Based on the target depth data corresponding to the reference image and the depth data corresponding to the rendered image, calculate the depth loss value corresponding to each pixel in the reference image;

[0209] Based on the reference image, the rendered image, and the preset semantic model, we obtain the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image.

[0210] Based on the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image, calculate the semantic loss value corresponding to each pixel in the reference image.

[0211] The Gaussian graph generation device for vehicle scenes provided in this embodiment is used to execute the technical solutions in any of the aforementioned method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0212] Figure 5 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 5 As shown, the electronic device 50 includes:

[0213] Processor 51, memory 52, and communication interface 53;

[0214] Memory 52 is used to store executable instructions of processor 51;

[0215] The processor 51 is configured to execute the technical solutions in any of the foregoing method embodiments by executing executable instructions.

[0216] Optionally, the memory 52 can be either standalone or integrated with the processor 51.

[0217] Optionally, when the memory 52 is a device independent of the processor 51, the electronic device 50 may further include:

[0218] Bus 54, memory 52 and communication interface 53 are connected to processor 51 through bus 54 and complete communication with each other. Communication interface 53 is used to communicate with other devices.

[0219] Optionally, the communication interface 53 can be implemented using a transceiver. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write databases, and read-only databases). The memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk drive.

[0220] Bus 54 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0221] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0222] The electronic device is used to execute the technical solutions in any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0223] This application also provides a vehicle, which includes a controller.

[0224] The controller is used to execute the technical solutions in any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0225] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, implements the technical solutions provided in any of the foregoing method embodiments.

[0226] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the technical solutions provided in any of the foregoing method embodiments.

[0227] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0228] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for generating Gaussian graphs of vehicle scenes, characterized in that, include: Gaussian graph generation data is obtained, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrix of each camera and the vehicle, and initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched. Based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined; Based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, determine the initial 3D Gaussian ellipsoid; Based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, the initial 3D Gaussian ellipsoid is optimized to generate a Gaussian map of the vehicle scene.

2. The method according to claim 1, characterized in that, The step of determining the target depth data and target camera pose data corresponding to each reference image based on the Gaussian map generation data and the multi-camera dense SLAM algorithm includes: Based on the Gaussian plot generation data and the multi-camera dense SLAM algorithm, the estimated depth data corresponding to each reference image is determined, as well as the reprojection loss value and depth loss value corresponding to each pixel in each reference image. Based on the reprojection loss value and depth loss value corresponding to each pixel in each reference image, and the multi-camera dense SLAM algorithm, the target optical flow coordinates and weights corresponding to each pixel in each reference image are determined. Based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image, a reprojection loss function and a depth loss function are generated. Based on the reprojection loss function, the depth loss function, and the pose transformation matrix corresponding to each camera and the vehicle, a target loss function is generated. Based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm, new initial depth data and new initial camera pose data are generated for each image in each image pair. Update the first iteration count; If the updated first iteration number is less than the first preset iteration threshold, repeat the above steps until the updated first iteration number is equal to the first preset iteration threshold. Then, use the new initial target depth data and the new initial camera pose data corresponding to each reference image as the target depth data and target camera pose data corresponding to each reference image.

3. The method according to claim 2, characterized in that, The step of determining the estimated depth data corresponding to each reference image, and the reprojection loss value and depth loss value corresponding to each pixel in each reference image, based on the Gaussian map generation data and the multi-camera dense SLAM algorithm, includes: Based on each image pair, the initial depth data and initial camera pose data corresponding to each reference image, the initial camera pose data corresponding to each image to be matched, the intrinsic parameters of each camera, and the multi-camera dense SLAM algorithm, the reprojection loss value of each pixel in each reference image is determined. Based on each reference image, the preset depth estimation model, and the initial depth data corresponding to each reference image, the estimated depth data corresponding to each reference image and the depth loss value corresponding to each pixel in each reference image are determined.

4. The method according to claim 2, characterized in that, The process of generating a reprojection loss function and a depth loss function based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, the intrinsic parameters of each camera, and the estimated depth data corresponding to each reference image includes: The reprojection loss function is generated based on the target optical flow coordinates and weights corresponding to each pixel in each reference image, as well as the intrinsic parameters of each camera. The depth loss function is generated based on the estimated depth data corresponding to each reference image.

5. The method according to claim 2, characterized in that, The process of generating new initial depth data and new initial camera pose data for each image in each image pair based on the target loss function, the pose transformation matrix corresponding to each camera and the vehicle, and the multi-camera dense SLAM algorithm includes: The vehicle pose data and new initial depth data corresponding to each image in each image pair are generated based on the target loss function and the multi-camera dense SLAM algorithm. For each reference image, based on the vehicle pose data and the pose transformation matrix between the camera that captured the reference image and the vehicle, new initial camera pose data corresponding to the reference image is obtained. Based on the new initial camera pose data corresponding to each reference image, generate new initial camera pose data corresponding to each image to be matched.

6. The method according to any one of claims 1 to 5, characterized in that, The step of optimizing the initial 3D Gaussian ellipsoid based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, to generate a Gaussian map of the vehicle scene includes: For each reference image, a rendering process is performed based on the target camera pose data corresponding to the reference image and the initial 3D Gaussian ellipsoid to obtain the rendered image corresponding to the reference image; For each reference image, based on the reference image, the target depth data corresponding to the reference image, and the rendered image, generate the loss value, semantic loss value, and depth loss value corresponding to each pixel in the reference image; The target loss value is calculated based on the photometric loss value, semantic loss value, and depth loss value corresponding to each pixel in each reference image. Based on the target loss value and the 3D Gaussian splashing technique, the initial 3D Gaussian ellipsoid is updated to obtain a new 3D Gaussian ellipsoid. Update the second iteration count; If the updated second iteration number is less than the second preset iteration threshold, repeat the above steps until the updated second iteration number is equal to the second preset iteration threshold, and use the image composed of the new 3D Gaussian ellipsoid as the Gaussian map of the vehicle scene.

7. The method according to claim 6, characterized in that, The step of generating photometric loss value, semantic loss value, and depth loss value for each pixel in the reference image based on the reference image, the target depth data corresponding to the reference image, and the rendered image includes: Based on the color data of each pixel in the reference image and the rendered image, calculate the photometric loss value corresponding to each pixel in the reference image; Based on the rendered image and the preset depth estimation model, generate depth data corresponding to the rendered image; Based on the target depth data corresponding to the reference image and the depth data corresponding to the rendered image, calculate the depth loss value corresponding to each pixel in the reference image; Based on the reference image, the rendered image, and the preset semantic model, the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image are obtained. Based on the semantic data corresponding to the reference image and the semantic data corresponding to the rendered image, the semantic loss value corresponding to each pixel in the reference image is calculated.

8. A Gaussian graph generation device for a vehicle scene, characterized in that, include: The acquisition module is used to acquire Gaussian map generation data, which includes multiple image pairs captured by multiple cameras in the vehicle, intrinsic parameters of each camera, pose transformation matrix of each camera and the vehicle, and initial depth data and initial camera pose data corresponding to each image in each image pair. Each image pair includes a reference image of the vehicle scene and an image to be matched. Processing module, used for: Based on the Gaussian graph generation data and the multi-camera dense SLAM algorithm, the target depth data and target camera pose data corresponding to each reference image are determined; Based on each reference image, and the target depth data and target camera pose data corresponding to each reference image, determine the initial 3D Gaussian ellipsoid; The generation module is used to optimize the initial 3D Gaussian ellipsoid based on each reference image, as well as the target depth data and target camera pose data corresponding to each reference image, to generate a Gaussian map of the vehicle scene.

9. An electronic device, characterized in that, include: Processor, memory, communication interface; The memory is used to store the executable instructions of the processor; The processor is configured to execute the Gaussian graph generation method for a vehicle scene according to any one of claims 1 to 7 by executing the executable instructions.

10. A vehicle, characterized in that, Including the controller; The controller is used to execute the Gaussian graph generation method for the vehicle scene as described in any one of claims 1 to 7.

11. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the Gaussian graph generation method for vehicle scenes as described in any one of claims 1 to 7.

12. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, is used to implement the Gaussian graph generation method for a vehicle scene as described in any one of claims 1 to 7.