Three-dimensional scene construction method and device in automatic driving environment, equipment and medium
By combining lidar and camera data in an autonomous driving environment, static and dynamic scene models are built and integrated, the problem of low reconstruction accuracy under single sensor modeling is solved, and three-dimensional scene reconstruction with higher accuracy is achieved, which improves the safety and efficiency of the autonomous driving system.
Patent Information
- Application Number
- CN202510369907.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, three-dimensional reconstruction has low reconstruction accuracy in autonomous driving environments, mainly based on single sensor modeling, resulting in poor scene recognition and tracking effects.
The multi-sensor data fusion method is adopted, combined with lidar point cloud data and camera image data, and through data preprocessing, dynamic and static scene targets, static and dynamic scene models are built, and model fusion is carried out to improve reconstruction accuracy.
Through the integration of multi-sensor data, the accuracy of three-dimensional scene reconstruction in an autonomous driving environment is effectively improved, the accuracy of obstacle identification and path planning is ensured, and the safety and efficiency of the autonomous driving system are improved.
Smart Images

Figure CN120374836A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of scene reconstruction, and more specifically, to a method, apparatus, device, and medium for constructing a three-dimensional scene in an autonomous driving environment. Background Art
[0002] Through three-dimensional reconstruction technology, an autonomous driving system can effectively detect and track obstacles on the road, which is crucial for ensuring driving safety and formulating effective path planning. The system can identify obstacles of various shapes and sizes and update their positions and states in real time to avoid collisions.
[0003] In related technologies, three-dimensional reconstruction is mainly achieved based on a single sensor model. For example, for a simple indoor scene, the point cloud data scanned by a lidar is directly voxelized, dividing the space into small cube units (voxels), and determining the approximate geometric structure of the scene according to the distribution of the point cloud in the voxels.
[0004] However, with the existing methods, the reconstruction accuracy is not high. Summary of the Invention
[0005] Embodiments described herein provide a method, apparatus, device, and medium for constructing a three-dimensional scene in an autonomous driving environment, which overcome the above problems.
[0006] In a first aspect, according to the content of the present disclosure, a method for constructing a three-dimensional scene in an autonomous driving environment is provided, including:
[0007] Obtaining current sensor data collected from the autonomous driving environment, where the current sensor data includes: first point cloud data and first image data;
[0008] Respectively performing data preprocessing on the first point cloud data and the first image data to obtain sensor processed data, where the sensor processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system;
[0009] Identifying corresponding dynamic scene targets and static scene targets in the autonomous driving environment from the sensor processed data;
[0010] Based on the second point cloud data and the second image data, constructing a static scene model corresponding to the static scene target;
[0011] Based on the second point cloud data, the second image data, and pre-collected IMU data, constructing a dynamic scene model corresponding to the dynamic scene target;
[0012] Fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
[0013] In a second aspect, according to the content of the present disclosure, there is provided a three-dimensional scene construction device in an autonomous driving environment, including:
[0014] An acquisition module, configured to acquire current sensor data collected from the autonomous driving environment, where the current sensor data includes: first point cloud data and first image data;
[0015] A preprocessing module, configured to perform data preprocessing on the first point cloud data and the first image data respectively to obtain sensor processed data, where the sensor processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system;
[0016] An identification module, configured to identify dynamic scene targets and static scene targets corresponding to the autonomous driving environment from the sensor processed data;
[0017] A first construction module, configured to construct a static scene model corresponding to the static scene target based on the second point cloud data and the second image data;
[0018] A second construction module, configured to construct a dynamic scene model corresponding to the dynamic scene target based on the second point cloud data, the second image data, and pre-collected IMU data;
[0019] A fusion module, configured to fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
[0020] In a third aspect, there is provided a computer device, including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the steps of the three-dimensional scene construction method in the autonomous driving environment in any one of the above embodiments are implemented.
[0021] In a fourth aspect, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the three-dimensional scene construction method in the autonomous driving environment in any one of the above embodiments are implemented.
[0022] The 3D scene construction method in the autonomous driving environment provided by the embodiments of the present application obtains the current sensor data collected from the autonomous driving environment, where the current sensor data includes: first point cloud data and first image data; respectively preprocess the first point cloud data and the first image data to obtain sensor processed data, and the sensor processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system; identify the corresponding dynamic scene targets and static scene targets in the autonomous driving environment from the sensor processed data; construct a static scene model corresponding to the static scene targets based on the second point cloud data and the second image data; construct a dynamic scene model corresponding to the dynamic scene targets based on the second point cloud data, the second image data, and the pre-collected IMU data; fuse the static scene model corresponding to the static scene targets and the dynamic scene model corresponding to the dynamic scene targets to construct a 3D scene model corresponding to the autonomous driving environment. In this way, by separately constructing the static scene model and the dynamic scene model from multi-sensor data and performing model fusion, the reconstruction accuracy of the 3D scene is effectively improved.
[0023] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the embodiments of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be understood that the following described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where:
[0025] Figure 1 is a flowchart of a 3D scene construction method in the autonomous driving environment provided by the present disclosure.
[0026] Figure 2 is a structural diagram of a 3D scene construction device in the autonomous driving environment provided by the present disclosure.
[0027] Figure 3 is a structural diagram of a computer device provided by the present disclosure.
[0028] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts shall also fall within the scope of protection of the present disclosure.
[0030] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter of the present disclosure belongs. Further, it will be understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. As used herein, the statement of joining or coupling two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.
[0031] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase "embodiments" appearing in various places in the specification does not necessarily refer to the same embodiment, nor are they independent or alternative embodiments mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0032] As used herein, the term "and / or" is merely a description of an associative relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the existence of A, the simultaneous existence of A and B, and the existence of B. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after. Terms such as "first" and "second" are only used to distinguish one component (or a part of a component) from another component (or another part of a component).
[0033] In the description of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more (including two). Similarly, "a plurality of groups" refers to two or more groups (including two groups).
[0034] To enable those skilled in the art of this technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0035] Figure 1 is a schematic flowchart of a method for constructing a three-dimensional scene in an autonomous driving environment, as Figure 1As shown in the figure, the specific process of the three-dimensional scene construction method in the autonomous driving environment includes:
[0036] S110. Obtain the current sensor data collected from the autonomous driving environment. The current sensor data includes: the first point cloud data and the first image data.
[0037] Among them, a variety of sensor devices can be installed / connected on the driving vehicle to collect data in the autonomous driving environment. The first point cloud data is the lidar point cloud data collected by the lidar device for the autonomous driving environment. The first image data is the camera image data collected by the camera for the autonomous driving environment.
[0038] S120. Perform data preprocessing on the first point cloud data and the first image data respectively to obtain sensor processed data. The sensor processed data includes: the second point cloud data and the second image data.
[0039] Among them, the first point cloud data and the second image data correspond to the same coordinate system.
[0040] For the data preprocessing of the second point cloud data, statistical filtering and / or radius filtering can be used to remove the outlier points generated due to environmental interference or sensor errors to ensure the accuracy of the point cloud data; at the same time, perform data format conversion to convert the data format of the second point cloud data into a format suitable for subsequent processing and storage, such as the PCD format.
[0041] For the data preprocessing of the first image data, image enhancement processing can be performed, such as adjusting brightness, contrast, saturation, etc. to improve the image quality and facilitate subsequent image stitching and feature extraction; remove salt-and-pepper noise, Gaussian noise, etc. through image denoising to ensure the clarity of the image; and convert the image into a unified format such as JPEG or PNG, etc.
[0042] In some embodiments, performing data preprocessing on the first point cloud data and the first image data respectively to obtain sensor processed data includes:
[0043] Delete the outlier points in the first point cloud data to obtain the second point cloud data. The outlier points in the first point cloud data are determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data, and / or based on the number of laser points included in the preset surrounding circle corresponding to the first point cloud data.
[0044] Among them, the outlier points in the first point cloud data can be determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data, or based on the number of laser points included in the preset surrounding circle corresponding to the first point cloud data, or based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data and the number of laser points included in the preset surrounding circle corresponding to the first point cloud data.
[0045] For example, the outlier points in the first point cloud data can be determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data, which may include: calculating the average distance / standard deviation between each point and its neighborhood points in the first point cloud data; if there is a point whose average distance / standard deviation from its neighborhood points exceeds the preset threshold, then determining that point as an outlier point and deleting it from the first point cloud data.
[0046] Determined based on the number of laser points included in the preset surrounding circle corresponding to the first point cloud data, which may include: setting a sphere with a fixed radius (preset surrounding circle), and counting the number of points within the sphere; if the number is less than the preset value, then determining the points within the sphere as outlier points and deleting them.
[0047] Determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data and the number of laser points included in the preset surrounding circle corresponding to the first point cloud data, then first calculate the average distance / standard deviation between each point and its neighborhood points in the first point cloud data; if there is a point whose average distance / standard deviation from its neighborhood points exceeds the preset threshold, then determining that point as an outlier point and deleting it from the first point cloud data; afterwards, setting a sphere with a fixed radius (preset surrounding circle), and counting the number of points within the sphere; if the number is less than the preset value, then determining the points within the sphere as outlier points and deleting them.
[0048] Thus, it is convenient to effectively remove the abnormal points caused by environmental interference (such as strong reflection, occlusion, etc.) or sensor errors, and ensure the accuracy and consistency of the point cloud data.
[0049] Perform histogram equalization processing on the first image data, and perform image denoising on the first image data after histogram equalization processing to obtain the second image data.
[0050] Among them, histogram equalization enhances the contrast of the image by redistributing the gray value distribution of the image. For example, calculate the gray histogram of the first image data, and transform the gray value of each pixel according to the distribution of the gray histogram, so that the details in the originally darker or brighter regions of the first image data are clearer.
[0051] For image denoising, a bilateral filtering algorithm can be adopted, which takes into account both the spatial distance between pixels and the difference in pixel gray values. By setting the filtering parameters in the spatial domain and gray domain, such as the standard deviation in the spatial domain being 1 - 2 pixels and the standard deviation in the gray domain being 10 - 20 gray values, filtering is performed on the first image data after histogram equalization processing, which can not only remove salt-and-pepper noise, Gaussian noise, etc., but also better retain the edges and detailed information of the image.
[0052] Align the coordinate systems of the second point cloud data and the second image data.
[0053] Among them, unify the lidar coordinate system, camera coordinate system, and vehicle coordinate system, so that data from different sensors can be fused and processed in the same reference system.
[0054] For coordinate system alignment, a quaternion-based method can be adopted. For example, obtain the initial transformation relationship between the lidar coordinate system, camera coordinate system, and vehicle coordinate system through a calibration experiment, represent the rotation relationship using quaternions, and convert the corresponding collected data from its own coordinate system to the vehicle coordinate system, which is convenient for ensuring accurate fusion of data from different sensors in the same reference system.
[0055] S130. Identify the corresponding dynamic scene targets and static scene targets in the autonomous driving environment from the sensor processing data.
[0056] Among them, dynamic scene targets can be pedestrians, cyclists, vehicles, etc. Static scene targets can be roads, lane lines, traffic signs, traffic lights, etc.
[0057] Data annotation can be pre-performed on the relevant attributes of dynamic scene targets / static scene targets for subsequent identification. For example, the behavior actions of pedestrians (such as walking direction, cycling speed, vehicle turning) and key feature points (such as joint points, vehicle contour points), road types, lane line attributes (dashed line, solid line, color), traffic sign meanings, traffic light control logics, etc.
[0058] S140. Based on the second point cloud data and the second image data, construct a static scene model corresponding to the static scene target.
[0059] In some embodiments, constructing a static scene model corresponding to the static scene target based on the second point cloud data and the second image data includes:
[0060] Perform point cloud registration on the second point cloud data to obtain a three-dimensional geometric model of the static scene target.
[0061] Among them, the three-dimensional geometric models of static scenes such as roads and buildings can be reconstructed by using algorithms such as point cloud registration and triangulation on the second point cloud data.
[0062] Taking the ICP (Iterative Closest Point) algorithm as an example, the construction of a three-dimensional geometric model is described. In the second point cloud data, points with similar distances and geometric features are searched to determine a set of initial corresponding point pairs, and the transformation matrix (including rotation and translation) between these corresponding point pairs is calculated to minimize the distance between the point clouds; through multiple iterations, the corresponding point pairs and the transformation matrix are continuously updated until a certain convergence condition is met, such as the average distance between the point clouds being less than a set threshold (such as 0.01 meters).
[0063] Perform Poisson surface reconstruction on the three-dimensional geometric model of the static scene target to reconstruct the three-dimensional structure model of the static scene target.
[0064] Among them, the Poisson surface reconstruction algorithm can be used to generate a geometric structure model with smooth surface and accurate structure, showing the basic shape and spatial layout of the static scene.
[0065] For example, the point cloud data in the three-dimensional geometric model of the static scene target is represented as an indicator function, and the discrete Poisson equation is used to solve the gradient field of the function, and then the reconstructed surface model is obtained.
[0066] Perform multi-view analysis on the second image data to obtain the surface texture model of the static scene target.
[0067] Among them, the multi-view stereo (MVS) technology can be adopted. By analyzing the corresponding relationship between images from different perspectives, the surface texture of the static scene target can be restored. At the same time, combined with three-dimensional reconstruction methods in deep learning, such as depth estimation and voxel reconstruction algorithms based on convolutional neural networks (CNNs), complex object models can be reconstructed more efficiently and accurately, improving the reconstruction quality and efficiency.
[0068] For example, feature points are extracted from images from different perspectives through feature extraction algorithms (such as SIFT, SURF, etc.), and corresponding feature point pairs are found using feature matching algorithms; according to the principle of triangulation, combined with the internal and external parameters of the camera (obtained through calibration), the three-dimensional coordinates of the feature points are calculated, thereby restoring the surface texture of the object.
[0069] For 3D reconstruction methods in deep learning, taking depth estimation as an example, a convolutional neural network model is trained using a large amount of annotated image data and corresponding depth maps. During the training process, the network learns the mapping relationship between features such as the texture and color of the image and the depth information. In the inference stage, the input image is passed through the trained network to obtain the depth map, and then combined with a voxel reconstruction algorithm to convert the depth map into a 3D voxel model, further improving the accuracy and efficiency of the reconstruction.
[0070] Map the surface texture model of the static scene target to the 3D structure model of the static scene target to obtain the static scene model corresponding to the static scene target.
[0071] Among them, through texture mapping technology, the surface texture model is accurately fitted to the geometric surface of the corresponding 3D structure model of the static scene target, generating a static scene digital asset that has both an accurate spatial structure and a realistic visual effect, that is, the static scene model.
[0072] For example, preprocess the data in the 3D structure model and surface texture model of the static scene target, including operations such as coordinate normalization and resolution matching; use texture mapping technology to accurately fit the surface texture model to the corresponding geometric surface according to the surface parameterization of the 3D structure model. For example, by calculating the UV coordinates on the surface of the 3D structure model, map the pixel values of the points in the surface texture model to the corresponding UV coordinate positions to achieve accurate texture fitting.
[0073] S150. Based on the second point cloud data, the second image data, and the pre-collected IMU data, construct a dynamic scene model corresponding to the dynamic scene target.
[0074] In some embodiments, the IMU data is used to describe the motion posture of the dynamic scene target in the autonomous driving environment.
[0075] Based on the second point cloud data, the second image data, and the pre-collected IMU data, constructing a dynamic scene model corresponding to the dynamic scene target includes:
[0076] Perform point cloud registration on the second point cloud data to obtain the 3D geometric model of the dynamic scene target.
[0077] Among them, the second point cloud data can be used to reconstruct the 3D geometric models of dynamic scenes such as pedestrians and vehicles through algorithms such as point cloud registration and triangulation.
[0078] Perform Poisson surface reconstruction on the 3D geometric model of the dynamic scene target to reconstruct the 3D structure model of the dynamic scene target.
[0079] Among them, the Poisson surface reconstruction algorithm can be used to generate a geometric structure model with a smooth surface and accurate structure, showing the basic shape and spatial layout of the dynamic scene.
[0080] For example, the point cloud data in the three-dimensional geometric model of the dynamic scene target is represented as an indicator function, and the gradient field of the function is solved by using the discretized Poisson equation, and then the reconstructed surface model is obtained.
[0081] Perform multi-view analysis on the second image data to obtain the surface texture model of the dynamic scene target.
[0082] Among them, the multi-view stereo (MVS) technology can be adopted to restore the surface texture of the dynamic scene target by analyzing the corresponding relationship between images from different perspectives.
[0083] Based on the motion posture of the dynamic scene target in the autonomous driving environment, combine the surface texture model of the dynamic scene target and the three-dimensional structure model of the dynamic scene target to obtain the dynamic scene model corresponding to the dynamic scene target.
[0084] Among them, data preprocessing and alignment are performed, combined with motion capture data (i.e., IMU data), and the motion capture data includes: information such as the motion trajectory and posture change of pedestrians or vehicles obtained through specialized motion capture devices (such as optical motion capture systems, inertial motion capture devices, etc.). Using this information, the surface texture model and the three-dimensional structure model of the dynamic scene target are combined and updated according to their actual motion states to construct a dynamic digital asset model with real motion characteristics, that is, the dynamic scene model.
[0085] In some embodiments, based on the motion posture of the dynamic scene target in the autonomous driving environment, combine the surface texture model of the dynamic scene target and the three-dimensional structure model of the dynamic scene target to obtain the dynamic scene model corresponding to the dynamic scene target, including:
[0086] Based on the motion posture of the dynamic scene target in the autonomous driving environment, obtain the posture-associated texture data from the surface texture model of the dynamic scene target; based on the motion posture of the dynamic scene target in the autonomous driving environment, obtain the posture-associated structure data from the three-dimensional structure model of the dynamic scene target; combine and update the posture-associated texture data and the posture-associated structure data to obtain the dynamic scene model corresponding to the dynamic scene target.
[0087] For example, the motion posture of a dynamic scene target in an autonomous driving environment is that the vehicle turns right. The posture-associated texture data is the texture data corresponding to the vehicle turning right process in the surface texture model of the dynamic scene target, and the posture-associated structure data is the geometric structure data corresponding to the vehicle turning right in the three-dimensional structure model of the dynamic scene target, such as the steering angle and moving path of the vehicle. By combining and updating the posture-associated texture data and the posture-associated structure data, it is convenient to accurately obtain the dynamic scene model corresponding to the dynamic scene target.
[0088] In some embodiments, it further includes:
[0089] Perform Kalman filtering on the IMU data to filter the IMU data; based on the acceleration measurement value and angular velocity measurement value of the moving vehicle, perform optimal state estimation on the IMU data to calibrate the IMU data.
[0090] Among them, Kalman filtering algorithm can be used for IMU data filtering. Using the measurement model and system state equation of the IMU, combined with the measurement values of acceleration and angular velocity, perform optimal estimation on the state of the IMU (such as position, velocity, attitude, etc.) to eliminate high-frequency noise and drift error in the measurement process.
[0091] In the process of scene reconstruction, the second point cloud data provides the basic geometric structure information of the scene, the second image data is used to supplement texture details and identify object features, and the IMU data helps to determine the motion state and attitude change of the sensor, so as to better fuse multi-frame data and correct errors in scene reconstruction.
[0092] For example, when constructing a road model, the point cloud data determines the plane and edge geometry of the road, the image data provides the texture of the road surface (such as lane line color, road surface material, etc.), and the IMU data assists in judging the motion trajectory of the vehicle during the data collection process to ensure that the data collected at different times can be accurately stitched together.
[0093] S160. Fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
[0094] Among them, effective scenario simulation testing can be achieved by integrating static scenario digital assets and dynamic digital asset models. During the simulation testing process, dynamic traffic participants (i.e., static scenario targets) need to interact with the static scenario environment, such as vehicles driving on the road and pedestrians walking on the sidewalk. To achieve this interaction, the dynamic digital asset model is placed in the static scenario digital asset model, and collision detection and response processing are performed according to physical rules and traffic rules. For example, when a vehicle drives to the road boundary or collides with other objects, a collision event is detected through a collision detection algorithm (such as a collision detection method based on bounding boxes or distance fields). The collision force and response can be calculated according to the physical engine to update the motion states of the vehicle and related objects, ensuring the authenticity and rationality of the entire simulation scenario.
[0095] After the three-dimensional model reconstruction, the reconstructed digital asset model can be simplified through model simplification algorithms (such as vertex clustering, edge contraction, etc.). On the premise of ensuring the visual effect and simulation requirements, the number of faces and data volume of the model are reduced, and the rendering and simulation operation efficiency are improved. At the same time, check the integrity of the model's topological structure, repair possible geometric errors such as holes and self-intersections, and optimize the normal direction, material properties, etc. of the model to make it more in line with physical realism and rendering requirements.
[0096] In some embodiments, it further includes:
[0097] Group the vertices in the three-dimensional scene model based on their spatial positions. If the number of vertices in the same group is greater than or equal to the preset clustering number, update the position of each vertex based on the central position of the same group.
[0098] For example, group the vertices of the model according to their spatial positions, set a clustering number threshold, and for the vertices that reach the clustering number, calculate their average position or adopt other simplification strategies (such as selecting the central vertex) to represent this group of vertices, thereby reducing the number of vertices.
[0099] In addition, the vertices of the model can also be grouped according to their spatial positions, set a clustering radius, and for the vertices within the clustering radius, calculate their average position or adopt other simplification strategies (such as selecting the central vertex) to represent this group of vertices, thereby reducing the number of vertices.
[0100] Or; update the connecting faces and topological relationships of the connecting edges based on the associated edges of the connecting edges in the three-dimensional scene model to simplify the three-dimensional scene model; where the associated edge is one or more model edges whose distance from the connecting edge is less than or equal to the preset distance.
[0101] Among them, by merging two endpoints or multiple endpoints into one vertex and updating the topological relationships of the faces and edges connected to the edge. During the contraction process, by considering maintaining the rationality of the geometric shape and topological structure of the model, problems such as self-intersection or excessive deformation are avoided.
[0102] Thus, on the premise of ensuring the visual effect and simulation requirements, the number of faces and the data volume of the model are significantly reduced, enabling the simulation test to run smoothly on ordinary computer hardware, reducing the calculation time and memory occupancy, and improving the test efficiency. At the same time, the simplified model is also more efficient in data transmission and storage, facilitating rapid data interaction and sharing between different modules and systems. In addition, appropriate model simplification can also reduce the noise and redundant information in the model, improve the quality and stability of the model, and is beneficial to subsequent algorithm testing and analysis.
[0103] In this embodiment, the current sensor data collected from the autonomous driving environment is obtained. The current sensor data includes: the first point cloud data and the first image data; the first point cloud data and the first image data are respectively preprocessed to obtain the sensor processed data. The sensor processed data includes: the second point cloud data and the second image data, and the second point cloud data and the second image data correspond to the same coordinate system; the dynamic scene targets and static scene targets corresponding to the autonomous driving environment are identified from the sensor processed data; based on the second point cloud data and the second image data, a static scene model corresponding to the static scene target is constructed; based on the second point cloud data, the second image data and the pre-collected IMU data, a dynamic scene model corresponding to the dynamic scene target is constructed; the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target are fused to construct a three-dimensional scene model corresponding to the autonomous driving environment. In this way, by separately constructing the static scene model and the dynamic scene model from multi-sensor data and performing model fusion, the reconstruction accuracy of the three-dimensional scene is effectively improved.
[0104] Figure 2 FIG. is a schematic structural diagram of a three-dimensional scene construction device in an autonomous driving environment provided in this embodiment. The three-dimensional scene construction device in the autonomous driving environment may include: an acquisition module 210, a preprocessing module 220, an identification module 230, a first construction module 240, a second construction module 250, and a fusion module 260.
[0105] The acquisition module 210 is configured to obtain the current sensor data collected from the autonomous driving environment. The current sensor data includes: the first point cloud data and the first image data.
[0106] A preprocessing module 220, configured to perform data preprocessing on the first point cloud data and the first image data respectively to obtain sensor-processed data, where the sensor-processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system.
[0107] An identification module 230, configured to identify corresponding dynamic scene targets and static scene targets in the autonomous driving environment from the sensor-processed data.
[0108] A first construction module 240, configured to construct a static scene model corresponding to the static scene target based on the second point cloud data and the second image data.
[0109] A second construction module 250, configured to construct a dynamic scene model corresponding to the dynamic scene target based on the second point cloud data, the second image data, and pre-collected IMU data.
[0110] A fusion module 260, configured to fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
[0111] In this embodiment, optionally, the first construction module 240 is specifically configured to:
[0112] Perform point cloud registration on the second point cloud data to obtain a three-dimensional geometric model of the static scene target; perform Poisson surface reconstruction on the three-dimensional geometric model of the static scene target to reconstruct a three-dimensional structure model of the static scene target; perform multi-view analysis on the second image data to obtain a surface texture model of the static scene target; map the surface texture model of the static scene target into the three-dimensional structure model of the static scene target to obtain a static scene model corresponding to the static scene target.
[0113] In this embodiment, optionally, the IMU data is used to describe the motion posture of the dynamic scene target in the autonomous driving environment.
[0114] The second construction module 250 is specifically configured to:
[0115] Perform point cloud registration on the second point cloud data to obtain a three-dimensional geometric model of the dynamic scene target; perform Poisson surface reconstruction on the three-dimensional geometric model of the dynamic scene target to reconstruct a three-dimensional structure model of the dynamic scene target; perform multi-view analysis on the second image data to obtain a surface texture model of the dynamic scene target; combine the surface texture model of the dynamic scene target and the three-dimensional structure model of the dynamic scene target based on the motion posture of the dynamic scene target in the autonomous driving environment to obtain a dynamic scene model corresponding to the dynamic scene target.
[0116] In this embodiment, optionally, the second construction module 250 is specifically configured to:
[0117] Based on the motion posture of the dynamic scene target in the autonomous driving environment, obtain pose-associated texture data from the surface texture model of the dynamic scene target; based on the motion posture of the dynamic scene target in the autonomous driving environment, obtain pose-associated structure data from the three-dimensional structure model of the dynamic scene target; combine and update the pose-associated texture data and the pose-associated structure data to obtain a dynamic scene model corresponding to the dynamic scene target.
[0118] In this embodiment, optionally, it further includes: a processing module and a calibration module.
[0119] The processing module is used to perform Kalman filtering on the IMU data to filter the IMU data.
[0120] The calibration module is used to perform optimal state estimation on the IMU data based on the acceleration measurement value and the angular velocity measurement value of the driving vehicle to calibrate the IMU data.
[0121] In this embodiment, optionally, the preprocessing module 220 is specifically used for:
[0122] Delete the outlier points in the first point cloud data to obtain the second point cloud data, where the outlier points in the first point cloud data are determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data, and / or determined based on the number of laser points included in the preset surrounding circle corresponding to the first point cloud data; perform histogram equalization processing on the first image data, and perform image denoising on the first image data after histogram equalization processing to obtain the second image data; align the coordinate systems of the second point cloud data and the second image data.
[0123] In this embodiment, optionally, it further includes: a simplification module.
[0124] The simplification module is used to group the vertices in the three-dimensional scene model based on the spatial position. If the number of vertices in the same group is greater than or equal to the preset clustering number, update the position of each vertex based on the central position of the same group; or; update the connection surface and topological relationship of the connection edge based on the associated edges of the connection edge in the three-dimensional scene model to simplify the three-dimensional scene model; where the associated edge is one or more model edges whose distance from the connection edge is less than or equal to the preset distance.
[0125] The three-dimensional scene construction device in the autonomous driving environment provided by the present disclosure can execute the above method embodiments, and its specific implementation principle and technical effects can be referred to the above method embodiments, which are not elaborated herein by the present disclosure.
[0126] The embodiment of the present application also provides a computer device. Specifically, please refer to Figure 3 ,Figure 3 This is the basic structural block diagram of the computer device in this embodiment.
[0127] The computer device includes a memory 310 and a processor 320 that are communicatively connected to each other through a system bus. It should be noted that only the computer device with the memory 310 and the processor 320 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of this technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0128] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through methods such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.
[0129] The memory 310 includes at least one type of readable storage medium, and the readable storage medium includes non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. The RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device. Of course, the memory 310 may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory 310 is generally used to store the operating system and various application software installed on the computer device, such as the program code of the above method. In addition, the memory 310 may also be used to temporarily store various data that have been output or will be output.
[0130] The processor 320 is generally used to execute the overall operations of the computer device. In this embodiment, the memory 310 is used to store program code or instructions, and the program code includes computer operation instructions. The processor 320 is used to execute the program code or instructions stored in the memory 310 or process data, such as running the program code of the above method.
[0131] In this document, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0132] Another embodiment of this application also provides a computer-readable medium, which can be a computer-readable signal medium or a computer-readable medium. The processor in the computer reads the computer-readable program code stored in the computer-readable medium, enabling the processor to perform the functional actions specified in each step or the combination of steps in the above method; and generating a device for performing the functional actions specified in each block or the combination of blocks in the block diagram.
[0133] The computer-readable medium includes but is not limited to electronic, magnetic, optical, electromagnetic, infrared memories or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. The memory is used to store program code or instructions, and the program code includes computer operation instructions. The processor is used to execute the program code or instructions of the above method stored in the memory.
[0134] For the definitions of the memory and the processor, reference can be made to the description of the foregoing computer device embodiments, and details are not repeated here.
[0135] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.
[0136] In each embodiment of this application, each functional unit or module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0137] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0138] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The "including" described in this application does not exclude the existence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the existence of a plurality of such elements. This application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the claims enumerating several units of a device, several of these units of the device can be embodied by the same hardware item. The use of first, second, and third, etc. does not denote any order and these words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
[0139] The above embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.
Claims
1. A method for constructing a three-dimensional scene in an autonomous driving environment, characterized in that Including: Obtain current sensor data collected from the autonomous driving environment, where the current sensor data includes: first point cloud data and first image data; Perform data preprocessing on the first point cloud data and the first image data respectively to obtain sensor processed data, where the sensor processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system; Identify corresponding dynamic scene targets and static scene targets in the autonomous driving environment from the sensor processed data; Based on the second point cloud data and the second image data, construct a static scene model corresponding to the static scene target; Based on the second point cloud data, the second image data, and pre-collected IMU data, construct a dynamic scene model corresponding to the dynamic scene target; Fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
2. The method according to claim 1, wherein The constructing a static scene model corresponding to the static scene target based on the second point cloud data and the second image data includes: Perform point cloud registration on the second point cloud data to obtain a three-dimensional geometric model of the static scene target; Perform Poisson surface reconstruction on the three-dimensional geometric model of the static scene target to reconstruct a three-dimensional structure model of the static scene target; Perform multi-view analysis on the second image data to obtain a surface texture model of the static scene target; Map the surface texture model of the static scene target into the three-dimensional structure model of the static scene target to obtain a static scene model corresponding to the static scene target.
3. The method according to claim 1, characterized in that, The IMU data is used to describe the motion posture of the dynamic scene target in the autonomous driving environment; Based on the second point cloud data, the second image data, and pre-collected IMU data, the constructing a dynamic scene model corresponding to the dynamic scene target includes: Perform point cloud registration on the second point cloud data to obtain a three-dimensional geometric model of the dynamic scene target; Perform Poisson surface reconstruction on the three-dimensional geometric model of the dynamic scene target to reconstruct a three-dimensional structure model of the dynamic scene target; Perform multi-view analysis on the second image data to obtain a surface texture model of the dynamic scene target; Based on the motion posture of the dynamic scene target in the autonomous driving environment, combine the surface texture model of the dynamic scene target and the three-dimensional structure model of the dynamic scene target to obtain a dynamic scene model corresponding to the dynamic scene target.
4. The method according to claim 3, wherein The combining the surface texture model of the dynamic scene target and the three-dimensional structure model of the dynamic scene target based on the motion posture of the dynamic scene target in the autonomous driving environment to obtain a dynamic scene model corresponding to the dynamic scene target includes: Based on the motion posture of the dynamic scene target in the autonomous driving environment, obtain pose-associated texture data from the surface texture model of the dynamic scene target; Obtain pose-associated structure data from the three-dimensional structure model of the dynamic scene target based on the motion pose of the dynamic scene target in the autonomous driving environment; Combine and update the pose-associated texture data and the pose-associated structure data to obtain a dynamic scene model corresponding to the dynamic scene target.
5. The method according to claim 4, characterized in that Further includes: Perform Kalman filtering on the IMU data to filter the IMU data; Based on the acceleration measurement value and the angular velocity measurement value of the driving vehicle, perform optimal state estimation on the IMU data to calibrate the IMU data.
6. The method according to claim 1, characterized in that, The step of respectively performing data preprocessing on the first point cloud data and the first image data to obtain sensor-processed data includes: Delete the outlier points in the first point cloud data to obtain a second point cloud data, where the outlier points in the first point cloud data are determined based on the distance statistical value between each laser point and the corresponding neighborhood points in the first point cloud data, and / or determined based on the number of laser points included in the preset surrounding circle corresponding to the first point cloud data; Perform histogram equalization processing on the first image data, and perform image denoising on the first image data after histogram equalization processing to obtain a second image data; Align the coordinate systems of the second point cloud data and the second image data.
7. The method according to claim 1, characterized in that Further includes: Group the vertices in the three-dimensional scene model based on the spatial position. If the number of vertices in the same group is greater than or equal to the preset clustering number, update the position of each vertex based on the central position of the same group; Or; Update the connection surface and topological relationship of the connection edge based on the associated edges of the connection edge in the three-dimensional scene model to simplify the three-dimensional scene model; Wherein, the associated edge is one or more model edges whose distance from the connection edge is less than or equal to the preset distance.
8. A three-dimensional scene construction device in an autonomous driving environment, characterized in that Includes: An acquisition module, configured to acquire current sensor data collected from the autonomous driving environment, where the current sensor data includes: first point cloud data and first image data; A preprocessing module, configured to respectively perform data preprocessing on the first point cloud data and the first image data to obtain sensor-processed data, where the sensor-processed data includes: second point cloud data and second image data, and the second point cloud data and the second image data correspond to the same coordinate system; An identification module, configured to identify the corresponding dynamic scene target and static scene target in the autonomous driving environment from the sensor-processed data; A first construction module, configured to construct a static scene model corresponding to the static scene target based on the second point cloud data and the second image data; A second construction module, configured to construct a dynamic scene model corresponding to the dynamic scene target based on the second point cloud data, the second image data, and the pre-collected IMU data; A fusion module, configured to fuse the static scene model corresponding to the static scene target and the dynamic scene model corresponding to the dynamic scene target to construct a three-dimensional scene model corresponding to the autonomous driving environment.
9. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the three-dimensional scene construction method in the autonomous driving environment as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional scene construction method in the autonomous driving environment as described in any one of claims 1 to 7.