Vehicle driving scene reconstruction method and system, storage medium and equipment
By setting the rotation parameters and offset parameters in the Gaussian model for Gaussian ellipsoid correction, and combining the Euclidean distance and opacity constraint function, the model training noise problem caused by image jitter is solved, and the accuracy and stability of three-dimensional reconstruction are improved.
Patent Information
- Application Number
- CN202510094753.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
The existing 3D Gaussian-based three-dimensional reconstruction method can easily lead to model training noise when image jitters, which is difficult to effectively solve.
By setting the rotation parameters and offset parameters in the Gaussian model, the position and shape of the Gaussian ellipsoid are corrected to reduce the noise impact caused by jitter, and further optimize the model training results through the Euclidean distance constraint function and the opaque constraint function.
It effectively reduces the model training noise caused by image jitter, improves the accuracy and stability of three-dimensional reconstruction, and avoids the glitches of the suspended objects of the ground Gaussian ellipsoid and non-ground Gaussian ellipsoid.
Smart Images

Figure CN119941997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional reconstruction, and in particular to a method, system, storage medium and device for reconstructing a vehicle driving scene. Background Art
[0002] In recent years, with the rapid development of autonomous driving technology, more and more car companies, Internet and other technology companies have invested in the research and development of autonomous driving algorithms, among which the training and reliability testing of autonomous driving algorithms are particularly important. Due to the high cost of real vehicle testing and certain risks, it is difficult to simulate scenes such as pedestrians "looking out of the head". This requires data collection by modifying the data collection vehicle. The data for training autonomous driving algorithms needs to be collected and labeled on a large scale, which consumes a lot of manpower and time. Moreover, for corner cases such as pedestrians "looking out of the head", it is difficult for the data collection vehicle to collect similar data. However, by using the 3D Gaussian method to reconstruct the scene in three dimensions using the data of the data collection vehicle, not only can a large amount of training data be expanded and a large amount of human resources be saved, but the scene reconstructed by 3D Gaussian can also be used for scene editing in Unreal Engine, and scenes similar to pedestrians "looking out of the head" can be manually edited, which greatly enriches the training and testing scenes of autonomous driving.
[0003] The 3D Gaussian-based three-dimensional reconstruction method in the prior art usually obtains point cloud data through laser scanning, builds a Gaussian model based on the point cloud data, and inputs a large number of patterns into the Gaussian model for training the Gaussian model. However, in this method, if the camera shakes when the pattern is acquired, and the model is trained with these shaken patterns, it is easy to cause noise in the model training. Summary of the invention
[0004] In view of this, the present invention provides a method, system, storage medium and device for reconstructing a vehicle driving scene, aiming to reduce the noise impact caused by image jitter on model training.
[0005] In order to achieve the above-mentioned object, a first aspect of an embodiment of the present application provides a method for reconstructing a vehicle driving scene, the method comprising: According to the point cloud data in the current scene, determine the ground point cloud data and non-ground point cloud data; The ground point cloud data and the non-ground point cloud data are respectively input into the corresponding Gaussian models for training, and the corresponding Gaussian ellipsoids are constructed and corrected by the rotation parameters and offset parameters in the corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; Performing rendering and synthesis processing on the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthetic image; Determine the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthetic image, wherein the total loss function is obtained based on a Euclidean distance constraint function constraining a target ground Gaussian ellipsoid and an opacity constraint function constraining a target non-ground Gaussian ellipsoid; When the training result is qualified, the vehicle driving scene is reconstructed based on the collected point cloud data and each qualified trained Gaussian model.
[0006] Optionally, the corresponding Gaussian ellipsoids are corrected by rotation parameters and offset parameters in the corresponding Gaussian models to obtain a target ground Gaussian ellipsoid corresponding to the ground point cloud data and a target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data, including: Through a position correction algorithm, the position of the Gaussian ellipsoid corresponding to each Gaussian model is corrected to obtain a preliminary ground Gaussian ellipsoid and a preliminary non-ground Gaussian ellipsoid, wherein the position correction algorithm is determined by a rotation parameter and an offset parameter; The expression of the position correction algorithm is:
[0007] in, Refers to the center position of the Gaussian ellipsoid after position correction, refers to the rotation parameter, refers to the offset parameter, Refers to the center position of the Gaussian ellipsoid corresponding to the Gaussian model before position correction; The shape of the initially corrected ground Gaussian ellipsoid and the initially corrected non-ground Gaussian ellipsoid is corrected by a shape correction algorithm to obtain a target ground Gaussian ellipsoid and a target non-ground Gaussian ellipsoid, wherein the shape correction algorithm is determined by a rotation parameter; The expression of the shape correction algorithm is:
[0008] in, refers to the rotation quaternion of the shape-corrected Gaussian ellipsoid, Refers to the transformation function of quaternion to rotation matrix, Refers to the transformation function of the rotation matrix to the quaternion. Refers to the rotation quaternion of the Gaussian ellipsoid before shape correction.
[0009] Optionally, determine a Euclidean distance constraint function, including: Determine a Euclidean distance constraint function by using the center position of a target ground Gaussian ellipsoid and the center positions of a plurality of multiplied ground Gaussian ellipsoids, wherein the multiplied ground Gaussian ellipsoid is obtained by multiplying the target ground Gaussian ellipsoid; The expression of the Euclidean distance constraint function is:
[0010] in, refers to the Euclidean distance constraint function, Refers to the first of the multiple Gaussian ellipsoids to be proliferated i The center position of a Gaussian ellipsoid; Refers to the i The center position of the proliferated ground Gaussian ellipsoid proliferated by the Gaussian ellipsoid is n is the number of Gaussian ellipsoids to be populated.
[0011] Optionally, determining the opaque constraint function includes: Determining an opacity constraint function by the opacity of a plurality of proliferated non-ground Gaussian ellipsoids, wherein the proliferated non-ground Gaussian ellipsoids are obtained by proliferating a target non-ground Gaussian ellipsoid; The expression of the opaque constraint function is:
[0012] in, refers to the opaque constraint function, Refers to the first i The opacity of a multiplied non-ground Gaussian ellipsoid, m Refers to the number of proliferating non-ground Gaussian ellipsoids.
[0013] Optionally, determine a total loss function, including: Determine a total loss function according to the Euclidean distance constraint function, the opacity constraint function, the image loss function and the first preset constraint weight, and the second preset constraint weight; The expression of the total loss function is:
[0014] in, L is the total loss function, refers to the image loss function, refers to the first preset constraint weight, Refers to the second preset constraint weight.
[0015] Optionally, the method further comprises: When the training result is unqualified, the first preset constraint weight and the second preset constraint weight in the total loss function are increased until the training is qualified.
[0016] Optionally, rendering and synthesizing the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthesized image includes: Rendering the target ground Gaussian ellipsoid to obtain multiple ground images; Rendering the target non-ground Gaussian ellipsoid to obtain multiple non-ground images; The ground image and the non-ground image that correspond to each other are merged to obtain a composite image.
[0017] A second aspect of an embodiment of the present application provides a vehicle driving scene reconstruction system, the system comprising: A point cloud data determination module, used to determine ground point cloud data and non-ground point cloud data according to the point cloud data in the current scene; A target Gaussian ellipsoid determination module is used to input the ground point cloud data and the non-ground point cloud data into their respective corresponding Gaussian models for training, construct their respective corresponding Gaussian ellipsoids, and correct their respective corresponding Gaussian ellipsoids through rotation parameters and offset parameters in their respective corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; A synthetic image determination module is used to render and synthesize the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthetic image; A training result determination module, used to determine the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthetic image, wherein the total loss function is obtained based on a Euclidean distance constraint function constraining a target ground Gaussian ellipsoid and an opacity constraint function constraining a target non-ground Gaussian ellipsoid; The scene reconstruction module is used to reconstruct the vehicle driving scene based on the collected point cloud data and each qualified trained Gaussian model when the training result is qualified.
[0018] A third aspect of an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the method for reconstructing a vehicle driving scene as described in any one of the first aspects of the present application are implemented.
[0019] A fourth aspect of an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the method for reconstructing a vehicle driving scene as described in any one of the first aspects of the present application are implemented.
[0020] A vehicle driving scene reconstruction method provided by the present application is adopted, and the method includes: determining ground point cloud data and non-ground point cloud data according to point cloud data in the current scene; inputting the ground point cloud data and the non-ground point cloud data into their respective corresponding Gaussian models for training, constructing their respective corresponding Gaussian ellipsoids and correcting their respective corresponding Gaussian ellipsoids through rotation parameters and offset parameters in their respective corresponding Gaussian models to obtain a target ground Gaussian ellipsoid corresponding to the ground point cloud data and a target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; rendering and synthesizing the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthesized image; determining the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthesized image, the total loss function being obtained based on a Euclidean distance constraint function constraining the target ground Gaussian ellipsoid and an opaque constraint function constraining the target non-ground Gaussian ellipsoid; when the training result is qualified, reconstructing the vehicle driving scene based on the collected point cloud data and each qualified trained Gaussian model.
[0021] By acquiring the ground point cloud data and non-ground point cloud data of the current scene, the corresponding Gaussian models are trained to construct the ground Gaussian ellipsoid corresponding to the ground Gaussian model and the non-ground Gaussian ellipsoid corresponding to the non-ground Gaussian model, and then the Gaussian model is corrected based on the rotation parameters and offset parameters in the Gaussian model, that is, the position and shape of the Gaussian model are corrected to reduce the noise generated during training due to pattern jitter. The corrected target ground Gaussian ellipsoid and target non-ground Gaussian ellipsoid are then rendered and synthesized to obtain a synthetic image, and the synthetic image and the standard image taken in the current scene are input into the total loss function to determine the training result of the Gaussian model, and when the training is qualified, the vehicle driving scene is reconstructed based on the point cloud data and the Gaussian model. In addition, the Euclidean distance constraint function is added to the target ground Gaussian ellipsoid in the loss function to reduce the suspended matter in the target ground Gaussian ellipsoid, and the opacity constraint function is added to the target non-ground Gaussian ellipsoid in the loss function to avoid the burr phenomenon in the target non-ground Gaussian ellipsoid. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a step diagram of a vehicle driving scene reconstruction method proposed in an embodiment of the present application; Figure 2 is a flow chart of a method for reconstructing a vehicle driving scene provided by an embodiment of the present application; Figure 3 This is a flowchart of semantic segmentation mask processing provided by an embodiment of the present application; Figure 4 This is a scene reconstruction effect comparison diagram provided by an embodiment of the present application; Figure 5 This is a schematic diagram of point cloud distribution for ground scene reconstruction provided by an embodiment of the present application; Figure 6 It is a schematic diagram of a vehicle driving scene reconstruction system proposed in one embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0024] refer to Figure 1 and Figure 2 , Figure 1 is a step diagram of a vehicle driving scene reconstruction method proposed in an embodiment of the present application, Figure 2 FIG. 1 is a flow chart of a method for reconstructing a vehicle driving scene proposed in an embodiment of the present application. Figure 1 As shown, the method comprises the following steps: S11: Determine ground point cloud data and non-ground point cloud data according to the point cloud data in the current scene.
[0025] S12: Input the ground point cloud data and the non-ground point cloud data into their respective corresponding Gaussian models for training, construct their respective corresponding Gaussian ellipsoids, and correct their respective corresponding Gaussian ellipsoids through rotation parameters and offset parameters in their respective corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data.
[0026] In this embodiment, when laser radar scanning is used to obtain point cloud data of the vehicle driving to the current scene, the point cloud data is divided into ground point cloud data and non-ground point cloud data based on the point cloud positions recorded in the point cloud data. Point cloud data refers to a data form composed of a large number of points in the three-dimensional coordinate system of the current scene, and each point contains its coordinate information in three-dimensional space.
[0027] Since the ground in the scene usually has relatively flat, continuous and large-scale features, while the non-ground may contain various complex objects and terrain changes, in this embodiment, by modeling the ground and non-ground separately, more refined processing and optimization can be performed on the different features of the ground and non-ground, thereby improving the scene reconstruction accuracy. Specifically, this example is provided with a ground Gaussian model and a non-ground Gaussian model according to the point cloud data, and the ground point cloud data is input into the ground Gaussian model that has been set with various operating parameters to wait for the ground Gaussian model to be trained. After the ground point cloud data is input into the ground Gaussian model, the corresponding ground Gaussian ellipsoid is obtained, and the non-ground point cloud data is input into the non-ground Gaussian model that has been set with various operating parameters to wait for the non-ground Gaussian model to be trained. After the non-ground point cloud data is input into the non-ground Gaussian model, the corresponding non-ground Gaussian ellipsoid is obtained.
[0028] The collection vehicle obtains the standard pattern of the current scene by shooting with six cameras. However, during the collection process of the standard pattern, there may be objective reasons such as the camera is not fixed firmly and the collection vehicle is not driving smoothly, causing the image to be jittery. If the jittery standard pattern is input as a sample into the Gaussian model for model training, it will bring noise to the model training. In order to avoid the noise problem that may exist during model training, in this embodiment, a learnable rotation parameter and an offset parameter are set in the Gaussian model. Specifically, when the standard pattern is jittery, the non-ground elements and ground elements on the standard pattern are deviated from the standard pattern taken normally, and the standard pattern is used as the true value sample input during Gaussian model training. Therefore, in order to avoid the deviation of the Gaussian model training result due to the input of the deviated standard pattern, a learnable rotation parameter and an offset parameter are set in the Gaussian model. The rotation parameter and the offset parameter are used to correct the center position and shape of the Gaussian ellipsoid corresponding to the Gaussian model, and the Gaussian model with the corrected position and shape is trained with the standard pattern, thereby reducing the deviation caused by the jitter of the standard pattern. After correction, the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data are obtained.
[0029] S13: Rendering and synthesizing the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthesized image.
[0030] Specifically, since the Gaussian ellipsoid is a 3D sphere and the standard image used for the true value is a two-dimensional pattern, in this embodiment, the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid are rendered separately to obtain the two-dimensional image corresponding to each Gaussian ellipsoid, and then the corresponding ground image and non-ground image are synthesized to obtain the corresponding synthesized image, which is the image of a part of the current scene.
[0031] S14: Based on the total loss function, the standard image in the current scene collected, and the synthetic image, the training results of each Gaussian model are determined. The total loss function is obtained based on the Euclidean distance constraint function constraining the target ground Gaussian ellipsoid and the opaque constraint function constraining the target non-ground Gaussian ellipsoid. In model training, the loss function is used to measure the difference between the model prediction result and the actual observation value. In this embodiment, the model prediction result is the synthetic image, and the actual observation value is the standard image in the current scene collected. Substituting the standard image and the synthetic image into the loss function can determine the difference between the synthetic image and the standard image.
[0032] Specifically, the segment anything technique is used to perform semantic segmentation on the standard image to obtain the semantic information of ground elements and non-ground elements, and then the ground and non-ground masks are obtained according to the semantic information to divide the standard image into two parts, such as Figure 3 As shown, Figure 3 A semantic segmentation mask processing flow chart is provided for the present application, wherein the mask is used to indicate pixels or areas in a specific area of an image, so the ground part in the standard image and the ground part in the synthetic image can be determined by the mask value corresponding to the ground element, and then the ground loss function is determined according to the ground part of the standard image and the synthetic image. Since the proliferated ground Gaussian ellipsoid corresponding to the ground Gaussian model is proliferated during the training process, if the distance from the target ground Gaussian ellipsoid is too far, floating objects may be introduced in the process of rendering the ground, so this embodiment also constrains the target ground Gaussian ellipsoid corresponding to the ground Gaussian model, that is, adds a Euclidean distance constraint function, and specifically adds the Euclidean distance constraint function to the ground loss function to obtain the ground total loss function.
[0033] The non-ground part in the standard image and the non-ground part in the synthetic image can be determined by the mask value corresponding to the non-ground element, and then the non-ground loss function is determined according to the non-ground part of the standard image and the synthetic image. Since the target non-ground Gaussian ellipsoid corresponding to the non-ground Gaussian model is proliferated in the proliferated ground Gaussian ellipsoid, the semi-transparent non-ground Gaussian ellipsoid will accumulate pixel colors and cause burrs during the rendering process, the non-ground Gaussian ellipsoid corresponding to the non-ground Gaussian model is also constrained in this embodiment, that is, an opaque constraint function is added, and the opaque constraint function is specifically added to the non-ground loss function to obtain the non-ground total loss function.
[0034] The corresponding total loss function is obtained by adding the ground total loss function reflecting the training results of the ground Gaussian model and the non-ground total loss function reflecting the training results of the non-ground Gaussian model. The training results of the ground Gaussian model and the non-ground Gaussian model are determined by the total loss function. Specifically, if the value of the total loss function gradually decreases and tends to be stable, the training results of the ground Gaussian model and the non-ground Gaussian model are determined to be qualified. If the value of the total loss function continues to be high, the training results of the ground Gaussian model and the non-ground Gaussian model are determined to be unqualified.
[0035] S15: When the training result is qualified, the vehicle driving scene is reconstructed based on the collected point cloud data and each qualified trained Gaussian model.
[0036] Specifically, when the training result is qualified, the point cloud data in the collected scene is input into the qualified ground Gaussian model and the non-ground Gaussian model to reconstruct the scene, and the corresponding ground Gaussian ellipsoid and the non-ground Gaussian ellipsoid are obtained. The ground Gaussian ellipsoid and the non-ground Gaussian ellipsoid reflect the current driving scene of the vehicle. The present application also provides a scene reconstruction effect comparison diagram, such as Figure 4 As shown, Figure 4 Figure (a) refers to the scene reconstructed by the Gaussian ellipsoid without Euclidean distance constraint, and Figure (b) refers to the scene reconstructed by the Gaussian ellipsoid with Euclidean distance constraint. There are obvious floating objects in the lower left corner of Figure (a), while there are no obvious floating objects in the lower left corner of Figure (b). The present application also provides a schematic diagram of point cloud distribution for ground scene reconstruction, such as Figure 5 As shown in the figure, the point cloud distribution on the ground road is flat, and there are basically no redundant points, which shows that the method of reconstructing the scene in the present application has a significant effect on the constraint of the ground Gaussian model.
[0037] By setting the learnable rotation parameters and offset parameters in the vehicle driving scene reconstruction method provided in the present application, the noise effect on the model training caused by the standard image jitter caused by the jitter of the camera of the acquisition vehicle during the image acquisition process is offset; at the same time, by setting the Euclidean distance constraint function for the ground Gaussian model, the problem of suspended objects appearing in the target ground Gaussian ellipsoid during training is avoided, and by setting the opacity constraint function for the non-ground Gaussian model, the problem of burrs appearing in the target non-ground Gaussian ellipsoid during training is avoided.
[0038] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, step S12 includes steps S21 to S22: S21: performing position correction on the Gaussian ellipsoids corresponding to the respective Gaussian models through a position correction algorithm to obtain a preliminary corrected ground Gaussian ellipsoid and a preliminary corrected non-ground Gaussian ellipsoid, wherein the position correction algorithm is determined by a rotation parameter and an offset parameter; S22: The expression of the position correction algorithm is:
[0039] in, Refers to the center position of the Gaussian ellipsoid after position correction, refers to the rotation parameter, refers to the offset parameter, Refers to the center position of the Gaussian ellipsoid corresponding to the Gaussian model before position correction.
[0040] S23: performing shape correction on the initially corrected ground Gaussian ellipsoid and the initially corrected non-ground Gaussian ellipsoid by a shape correction algorithm to obtain a target ground Gaussian ellipsoid and a target non-ground Gaussian ellipsoid, wherein the shape correction algorithm is determined by the rotation parameter; S24: The expression of the shape correction algorithm is:
[0041] in, refers to the rotation quaternion of the shape-corrected Gaussian ellipsoid, Refers to the transformation function of quaternion to rotation matrix, Refers to the transformation function of the rotation matrix to the quaternion. Refers to the rotation quaternion of the Gaussian ellipsoid before shape correction.
[0042] In order to avoid the situation where noise occurs during model training due to image jitter, this embodiment determines the position correction algorithm that can correct the position of the Gaussian ellipsoid corresponding to the Gaussian model by setting the rotation parameters and offset parameters. The center position of the Gaussian ellipsoid is corrected by the position correction algorithm, which avoids the position offset of the ground part or non-ground part in the standard image caused by the jitter of the standard image, which brings noise to the model training. Specifically, the expression of the position correction algorithm is: ,in, Refers to the center position of the Gaussian ellipsoid after position correction, refers to the rotation parameter, refers to the offset parameter, Refers to the center position of the Gaussian ellipsoid corresponding to the Gaussian model before position correction. The Gaussian ellipsoid is corrected by calculating the center position of the Gaussian ellipsoid after position correction. Specifically, the ground Gaussian ellipsoid is corrected by calculating the center position of the ground Gaussian ellipsoid after correction to obtain a preliminary ground Gaussian ellipsoid; the non-ground Gaussian ellipsoid is corrected by calculating the center position of the non-ground Gaussian ellipsoid after correction to obtain a preliminary non-ground Gaussian ellipsoid.
[0043] In order to avoid the situation where noise occurs during model training due to image jitter, this embodiment determines the position correction algorithm that can correct the shape of the Gaussian ellipsoid corresponding to the Gaussian model by setting the rotation parameters. The shape of the Gaussian ellipsoid is corrected by the shape correction algorithm, which avoids the deformation of the shape of the ground part or non-ground part in the standard image due to the jitter of the standard image, which brings the influence of noise to the model training. Specifically, the expression for the shape correction algorithm is: ,in, refers to the rotation quaternion of the shape-corrected Gaussian ellipsoid, Refers to the transformation function of quaternion to rotation matrix, Refers to the transformation function of the rotation matrix to the quaternion. Refers to the rotation quaternion of the Gaussian ellipsoid before shape correction. Among them, quaternion and rotation matrix are both used to describe the rotation of the Gaussian ellipsoid in three-dimensional space, including the rotation angle of the Gaussian ellipsoid, etc. When calculating the new coordinates of the rotated Gaussian ellipsoid, the rotation quaternion of the Gaussian ellipsoid before shape correction is first converted into a rotation matrix for calculation, and then the rotation matrix is converted back into a rotation quaternion to obtain the rotation shape of the corrected Gaussian ellipsoid. Specifically, the shape of the initially repaired ground Gaussian ellipsoid is corrected by calculating the corrected rotation shape of the initially repaired ground Gaussian ellipsoid to obtain the target ground Gaussian ellipsoid; the shape of the initially repaired non-ground Gaussian ellipsoid is corrected by calculating the corrected rotation shape of the initially repaired non-ground Gaussian ellipsoid to obtain the target non-ground Gaussian ellipsoid.
[0044] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, determining the Euclidean distance constraint function includes steps S31 to S32: S31: determining a Euclidean distance constraint function according to the center position of a target ground Gaussian ellipsoid and the center positions of a plurality of multiplied ground Gaussian ellipsoids, wherein the multiplied ground Gaussian ellipsoid is obtained by multiplying the target ground Gaussian ellipsoid; S32: The expression of the Euclidean distance constraint function is:
[0045] in, refers to the Euclidean distance constraint function, Refers to the first of the multiple Gaussian ellipsoids to be proliferated i The center position of a Gaussian ellipsoid; Refers to the i The center position of the proliferated ground Gaussian ellipsoid proliferated by the Gaussian ellipsoid is n is the number of Gaussian ellipsoids to be populated.
[0046] Since the road on which the vehicle travels is flat, during the training process, if the Gaussian ellipsoid of the target ground Gaussian ellipsoid is too far away from the target ground Gaussian ellipsoid, floating objects may be introduced in the process of rendering the ground. Therefore, in this embodiment, the target ground Gaussian ellipsoid is constrained to reduce the generation of suspended objects. Specifically, in order to more accurately represent the point cloud data in the scene, the target ground Gaussian ellipsoid will be proliferated to obtain multiple new Gaussian ellipsoids, that is, the proliferated ground Gaussian ellipsoids, so that the proliferated ground Gaussian ellipsoid can fill the ground part of the entire scene as much as possible to achieve the effect of ground rendering.
[0047] Specifically, the Gaussian ellipsoid proliferation strategy is to proliferate the Gaussian ellipsoid according to the size of the position gradient, that is, when the target ground Gaussian ellipsoid is proliferated for the first time, a proliferated ground Gaussian ellipsoid will be proliferated, and the target ground Gaussian ellipsoid at this time is the Gaussian ellipsoid to be proliferated; in the next proliferation, the target ground Gaussian ellipsoid and the proliferated proliferated ground Gaussian ellipsoid are used as new Gaussian ellipsoids to be proliferated, and the Gaussian ellipsoids to be proliferated are proliferated to obtain their own proliferated proliferated ground Gaussian ellipsoids; in the next new proliferation, the newly proliferated proliferated ground Gaussian ellipsoid and the original Gaussian ellipsoid to be proliferated are used as new Gaussian ellipsoids to be proliferated for new proliferation. In this embodiment, by obtaining the center positions of multiple proliferated ground Gaussian ellipsoids and the center position of the target ground Gaussian ellipsoid, a Euclidean distance constraint function that can constrain the target ground Gaussian ellipsoid can be determined, and the expression of the Euclidean distance constraint function is: in, refers to the Euclidean distance constraint function, Refers to the first of the multiple Gaussian ellipsoids to be proliferated i The center position of a Gaussian ellipsoid; Refers to the i The center position of the proliferated ground Gaussian ellipsoid proliferated by the Gaussian ellipsoid is n is the number of Gaussian ellipsoids to be populated.
[0048] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, determining an opacity constraint function includes steps S41 to S42: S41: determining an opacity constraint function by using the opacity of a plurality of proliferated non-ground Gaussian ellipsoids, wherein the proliferated non-ground Gaussian ellipsoids are obtained by proliferating the target non-ground Gaussian ellipsoid; S42: The expression of the opaque constraint function is:
[0049] in, refers to the opaque constraint function, Refers to the first i The opacity of a multiplied non-ground Gaussian ellipsoid, m Refers to the number of proliferating non-ground Gaussian ellipsoids.
[0050] Specifically, since during the rendering process, the translucent target non-ground Gaussian ellipsoid and the proliferated non-ground Gaussian ellipsoid will accumulate pixel colors, thereby causing burrs, in order to avoid this situation, the target non-ground Gaussian ellipsoid is constrained in this embodiment. Specifically, in order to more accurately represent the point cloud data in the scene, the target non-ground Gaussian ellipsoid will be proliferated to obtain multiple new Gaussian ellipsoids, that is, the proliferated non-ground Gaussian ellipsoids, so that the proliferated non-ground Gaussian ellipsoid can fill the non-ground part of the entire scene as much as possible, achieving the effect of non-ground rendering. By obtaining the opacity of each of the multiple proliferated non-ground Gaussian ellipsoids, the opacity constraint function of the target non-ground Gaussian ellipsoid is determined, and the expression of the opacity constraint function is:
[0051] in, refers to the opaque constraint function, It refers to the opacity of the i-th proliferated non-ground Gaussian ellipsoid obtained by proliferating the target non-ground Gaussian ellipsoid, and m refers to the number of proliferated non-ground Gaussian ellipsoids. In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, determining the total loss function includes steps S51 to S52: S51: determining a total loss function according to the Euclidean distance constraint function, the opacity constraint function, the image loss function, the first preset constraint weight, and the second preset constraint weight; S52: The expression of the total loss function is:
[0052] in, L is the total loss function, refers to the image loss function, refers to the first preset constraint weight, Refers to the second preset constraint weight.
[0053] Since the ground loss function can determine the ground part in the standard image and the ground part in the synthetic image through the mask value corresponding to the ground element, and then determine it according to the ground part of the standard image and the synthetic image, and the non-ground loss function can determine the non-ground part in the standard image and the non-ground part in the synthetic image through the mask value corresponding to the non-ground element, and then determine it according to the non-ground part of the standard image and the synthetic image, in this embodiment, the ground loss function and the non-ground loss function are integrated together to form an image loss function to reflect the difference between the standard image and the synthetic image. Specifically, the expression for integrating the ground loss function and the non-ground loss function to form the image loss function is:
[0054] in, and is a hyperparameter, indicating the weight of the constraint. Represents the mask of ground and non-ground elements obtained by segmentanything, Represents the standard image in the current scene, Represents a composite image. Represents the calculation of image structure similarity function, N represents the number of images. In this image loss function, the mask It is a matrix containing different pixel values, that is, each pixel value in the mask corresponds to a pixel in the image, and according to the difference in these pixel values, it can be distinguished which pixels belong to ground elements and which pixels belong to non-ground elements.
[0055] Since the total loss function is the sum of the ground total loss function and the non-ground total loss function, and the ground total loss function includes the ground loss function and the Euclidean distance constraint function, and the non-ground total loss function includes the non-ground loss function and the opaque constraint function, in this embodiment, the total loss function is composed of the image loss function, the Euclidean distance constraint function and the opaque constraint function. Since setting weights for the constraint function can adjust the optimization direction of the model during training, in this implementation, a first preset constraint weight is set for the Euclidean distance constraint function, and a second preset constraint weight is set for the opaque constraint function. Specifically, the expression of the total loss function determined by the Euclidean distance constraint function, the opaque constraint function, the image loss function, the first preset constraint weight, and the second preset constraint weight is:
[0056] in, L is the total loss function, refers to the image loss function, refers to the first preset constraint weight, Refers to the second preset constraint weight.
[0057] In combination with the above embodiments, in one implementation, the present invention also provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, step S54 includes step S61: S61: When the training result is unqualified, increase the first preset constraint weight and the second preset constraint weight in the total loss function until the training is qualified.
[0058] Specifically, when the training result is unqualified, the Gaussian model needs to be further optimized. Specifically, by increasing the first preset constraint weight and the second preset constraint weight in the loss function, the Gaussian model is guided to pay more attention to the two loss terms of Euclidean distance constraint and opaque constraint. If one adjustment fails to make the training result qualified, you can try multiple iterative optimizations, that is, in each iteration, adjust the constraint weights and other parameters of the model. Through multiple iterative optimizations, you can gradually approach the optimal solution, so that the value of the total loss function gradually decreases and tends to be stable, so that the training is qualified.
[0059] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for reconstructing a vehicle driving scene. In the method for reconstructing a vehicle driving scene, step S13 includes steps S71 to S73: S71: Rendering the target ground Gaussian ellipsoid to obtain multiple ground images.
[0060] S72: Rendering the target non-ground Gaussian ellipsoid to obtain a plurality of non-ground images.
[0061] S73: merging the ground image and the non-ground image that correspond to each other to obtain a composite image.
[0062] Since the Gaussian ellipsoid is a 3D sphere, in order to match it with a two-dimensional standard image, this embodiment renders the Gaussian ellipsoid to obtain a two-dimensional image corresponding to the Gaussian ellipsoid. Specifically, the target ground Gaussian ellipsoid is rendered to obtain multiple ground images, and the target non-ground Gaussian ellipsoid is rendered to obtain multiple non-ground images. Since a road image includes both ground images and non-ground images, there is a corresponding relationship between the ground image and the non-ground image. The ground image and the non-ground image that are in a corresponding relationship are merged to obtain a complete image, that is, a composite image. After merging multiple non-ground images with their corresponding ground images, multiple composite images are obtained.
[0063] Based on the same inventive concept, an embodiment of the present application provides a vehicle driving scene reconstruction system, referring to Figure 6 , Figure 6 is a schematic diagram of a vehicle driving scene reconstruction system proposed in an embodiment of the present application, the system comprising: A point cloud data determination module, used to determine ground point cloud data and non-ground point cloud data according to the point cloud data in the current scene; A target Gaussian ellipsoid determination module is used to input the ground point cloud data and the non-ground point cloud data into their respective corresponding Gaussian models for training, construct their respective corresponding Gaussian ellipsoids, and correct their respective corresponding Gaussian ellipsoids through rotation parameters and offset parameters in their respective corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; A synthetic image determination module is used to render and synthesize the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthetic image; A training result determination module, used to determine the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthetic image, wherein the total loss function is obtained based on a Euclidean distance constraint function constraining a target ground Gaussian ellipsoid and an opacity constraint function constraining a target non-ground Gaussian ellipsoid; The scene reconstruction module is used to reconstruct the vehicle driving scene based on the collected point cloud data and each qualified trained Gaussian model when the training result is qualified.
[0064] Based on the same inventive concept, another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the method for reconstructing a vehicle driving scene as described in any of the above embodiments are implemented.
[0065] Based on the same inventive concept, another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the vehicle driving scene reconstruction method as described in any of the above embodiments.
[0066] This application sets rotation parameters and offset parameters in the Gaussian model to correct the Gaussian ellipsoid corresponding to the Gaussian model through position correction algorithm and shape correction algorithm, thus avoiding the noise problem caused by image jitter in model training. In addition, by applying the Euclidean distance constraint function to the ground Gaussian ellipsoid corresponding to the ground Gaussian model, the appearance of suspended objects in the ground Gaussian ellipsoid during training is avoided; by applying the opacity constraint function to the non-ground Gaussian ellipsoid corresponding to the non-ground Gaussian model, the appearance of burrs in the non-ground Gaussian ellipsoid during rendering is avoided.
[0067] Those skilled in the art will appreciate that embodiments of the present invention may provide methods, devices, electronic devices, storage media, or computer program products. Therefore, embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Moreover, embodiments of the present invention may take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0068] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0069] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0070] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0071] In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the previously and subsequently associated objects are in an "or" relationship.
[0072] The above is a detailed introduction to the vehicle driving scene reconstruction method, system, storage medium and device provided by the present application. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A method for reconstructing a vehicle driving scene, characterized in that: The method comprises: According to the point cloud data in the current scene, determine the ground point cloud data and non-ground point cloud data; The ground point cloud data and the non-ground point cloud data are respectively input into the corresponding Gaussian models for training, and the corresponding Gaussian ellipsoids are constructed and corrected by the rotation parameters and offset parameters in the corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; Performing rendering and synthesis processing on the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthetic image; Determine the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthetic image, wherein the total loss function is obtained based on a Euclidean distance constraint function constraining a target ground Gaussian ellipsoid and an opacity constraint function constraining a target non-ground Gaussian ellipsoid; When the training result is qualified, the vehicle driving scene is reconstructed based on the collected point cloud data and each qualified trained Gaussian model.
2. The method for reconstructing a vehicle driving scene according to claim 1, characterized in that: The corresponding Gaussian ellipsoids are corrected by the rotation parameters and offset parameters in the corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data, including: Through a position correction algorithm, the position of the Gaussian ellipsoid corresponding to each Gaussian model is corrected to obtain a preliminary ground Gaussian ellipsoid and a preliminary non-ground Gaussian ellipsoid, wherein the position correction algorithm is determined by a rotation parameter and an offset parameter; The expression of the position correction algorithm is: in, Refers to the center position of the Gaussian ellipsoid after position correction, refers to the rotation parameter, refers to the offset parameter, Refers to the center position of the Gaussian ellipsoid corresponding to the Gaussian model before position correction; The shape of the initially corrected ground Gaussian ellipsoid and the initially corrected non-ground Gaussian ellipsoid is corrected by a shape correction algorithm to obtain a target ground Gaussian ellipsoid and a target non-ground Gaussian ellipsoid, wherein the shape correction algorithm is determined by a rotation parameter; The expression of the shape correction algorithm is: in, refers to the rotation quaternion of the shape-corrected Gaussian ellipsoid, Refers to the transformation function of quaternion to rotation matrix, Refers to the transformation function of the rotation matrix to the quaternion. Refers to the rotation quaternion of the Gaussian ellipsoid before shape correction.
3. The method for reconstructing a vehicle driving scene according to claim 1, characterized in that: Determine the Euclidean distance constraint function, including: Determine a Euclidean distance constraint function by using the center position of a target ground Gaussian ellipsoid and the center positions of a plurality of multiplied ground Gaussian ellipsoids, wherein the multiplied ground Gaussian ellipsoid is obtained by multiplying the target ground Gaussian ellipsoid; The expression of the Euclidean distance constraint function is: in, refers to the Euclidean distance constraint function, Refers to the first of the multiple Gaussian ellipsoids to be proliferated i The center position of a Gaussian ellipsoid; Refers to the i The center position of the proliferated ground Gaussian ellipsoid proliferated by the Gaussian ellipsoid is n is the number of Gaussian ellipsoids to be populated.
4. The method for reconstructing a vehicle driving scene according to claim 1, characterized in that: Functions that determine opacity constraints include: Determining an opacity constraint function by the opacity of a plurality of proliferated non-ground Gaussian ellipsoids, wherein the proliferated non-ground Gaussian ellipsoids are obtained by proliferating a target non-ground Gaussian ellipsoid; The expression of the opaque constraint function is: in, refers to the opaque constraint function, Refers to the first i The opacity of a multiplied non-ground Gaussian ellipsoid, m Refers to the number of proliferating non-ground Gaussian ellipsoids.
5. The vehicle driving scene reconstruction method according to claim 4, characterized in that: Determine the total loss function, including: Determine a total loss function according to the Euclidean distance constraint function, the opacity constraint function, the image loss function and the first preset constraint weight, and the second preset constraint weight; The expression of the total loss function is: in, L is the total loss function, refers to the image loss function, refers to the first preset constraint weight, Refers to the second preset constraint weight.
6. The vehicle driving scene reconstruction method according to claim 5, characterized in that: The method further comprises: When the training result is unqualified, the first preset constraint weight and the second preset constraint weight in the total loss function are increased until the training is qualified.
7. The vehicle driving scene reconstruction method according to claim 5, characterized in that: The target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid are rendered and synthesized to obtain a synthetic image, including: Rendering the target ground Gaussian ellipsoid to obtain multiple ground images; Rendering the target non-ground Gaussian ellipsoid to obtain multiple non-ground images; The ground image and the non-ground image that correspond to each other are merged to obtain a composite image.
8. A vehicle driving scene reconstruction system, characterized in that: The system comprises: A point cloud data determination module, used to determine ground point cloud data and non-ground point cloud data according to the point cloud data in the current scene; A target Gaussian ellipsoid determination module is used to input the ground point cloud data and the non-ground point cloud data into their respective corresponding Gaussian models for training, construct their respective corresponding Gaussian ellipsoids, and correct their respective corresponding Gaussian ellipsoids through rotation parameters and offset parameters in their respective corresponding Gaussian models to obtain the target ground Gaussian ellipsoid corresponding to the ground point cloud data and the target non-ground Gaussian ellipsoid corresponding to the non-ground point cloud data; A synthetic image determination module is used to render and synthesize the target ground Gaussian ellipsoid and the target non-ground Gaussian ellipsoid to obtain a synthetic image; A training result determination module, used to determine the training results of each Gaussian model based on a total loss function, a standard image in the current scene collected, and the synthetic image, wherein the total loss function is obtained based on a Euclidean distance constraint function constraining a target ground Gaussian ellipsoid and an opacity constraint function constraining a target non-ground Gaussian ellipsoid; The scene reconstruction module is used to reconstruct the vehicle driving scene based on the collected point cloud data and each qualified trained Gaussian model when the training result is qualified.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the vehicle driving scene reconstruction method as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the vehicle driving scene reconstruction method according to any one of claims 1 to 7 are implemented.