Method and device for establishing vehicle-road cloud integrated collaborative awareness function unified model
By establishing a unified model of integrated vehicle-road and cloud-based collaborative perception function in the cloud, reconstructing 4D scenes and virtual data, and training inference models, the limitations of on-board perception algorithms in adapting to the needs of multiple models, multiple platforms and multi-functions are solved, and efficient perceptual data processing and autonomous driving support are achieved.
Patent Information
- Application Number
- CN202411941297.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-03
AI Technical Summary
Existing on-board perception algorithms have limitations in the adaptability and performance improvement of complex tasks of perception models, especially in the difficulty of adapting to different vehicle models, multi-platform and multi-functional needs. Due to the huge amount of data, it is difficult for vehicle-side inferencers to achieve this requirement.
By establishing a unified model of vehicle-road and cloud integrated collaborative perception function based on sensor information of vehicle-road perception system in the cloud, reconstructing 4D scenes and virtual data of environmentally-aware data, using virtual data to train and fine-tune the inference model, and finally transmitting the perception results to the vehicle.
The reconstruction of 4D scenes is realized by integrating multi-source data information, meeting the needs of "multi-vehicle models/multi-platforms/multi-functionality", enhancing the robustness of the NeRF model to data, optimizing the inference time, reducing the computing pressure of the vehicle-side inference machine, and supporting the planning and control decisions of downstream autonomous driving.
Smart Images

Figure CN120086965A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of autonomous driving perception, and particularly relates to a method and device for establishing a unified model for vehicle-road-cloud integrated collaborative perception functions. Background Art
[0002] The perception system of intelligent driving vehicles uses various sensors such as cameras, lidar, and millimeter-wave radars to simulate the visual function of humans, and captures the driving environment around the vehicle in real time, including static elements such as lane markings, traffic lights, and traffic signs, as well as information on moving obstacles such as other vehicles and pedestrians. This module is a core component of intelligent driving vehicles and provides necessary inputs for subsequent positioning, prediction, decision-making, path planning, and control. Therefore, the accuracy and reliability of the underlying algorithms are crucial.
[0003] In related technologies, most mass-produced in-vehicle perception algorithms adopt neural network technologies based on deep learning, that is, deep neural networks are trained with a large amount of labeled data and applied to various visual perception tasks.
[0004] However, there are still some limitations in the adaptability of the perception models of in-vehicle perception algorithms in related technologies to complex tasks and performance improvement. For example, the position distributions of sensors and roadside devices on different vehicle models are relatively irregular, resulting in a model being difficult to meet the requirements of "multiple vehicle models / multiple platforms / multiple functions". Moreover, due to the huge amount of fused data, it is also difficult for in-vehicle inference engines to meet this requirement, which urgently needs to be solved. Summary of the Invention
[0005] This application provides a method, device, electronic device, and storage medium for establishing a unified model for vehicle-road-cloud integrated collaborative perception functions to solve the problems that there are still some limitations in the adaptability of the in-vehicle perception algorithms in related technologies to complex tasks and performance improvement. For example, the position distributions of sensors and roadside devices on different vehicle models are relatively irregular, resulting in a model being difficult to meet the requirements of "multiple vehicle models / multiple platforms / multiple functions". Moreover, due to the huge amount of fused data, it is also difficult for in-vehicle inference engines to meet this requirement, etc.
[0006] The first aspect embodiment of this application provides a method for establishing a unified model for vehicle-road-cloud integrated collaborative perception function, which is applied to the cloud. Among them, the method includes the following steps: obtaining sensor information of the perception system of the vehicle, and establishing a target combination model and a virtual camera of the perception system based on the sensor information and the requirements of the perception system; obtaining the environmental perception data of the perception system, using the target combination model to transmit the environmental perception data to the target platform, and reconstructing the 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data and the target NeRF model, so as to generate and output virtual data corresponding to the 4D scene by using the virtual camera; collecting the virtual data to train and fine-tune the target inference model to obtain the final inference model, so as to use the final inference model to infer the perception result of the environmental perception data, and transmitting the perception result to the vehicle to complete the establishment of the unified model for vehicle-road-cloud integrated collaborative perception function.
[0007] Optionally, in an embodiment of this application, the establishing the target combination model and the virtual camera of the perception system based on the sensor information and the requirements of the perception system includes: collecting the processing module information corresponding to the sensor information to establish the target combination model according to the processing module information; constructing a perception scheme according to the sensor information to establish the virtual camera through the perception scheme.
[0008] Optionally, in an embodiment of this application, before using the target combination model to transmit the environmental perception data to the target platform, it further includes: obtaining the data requirements and data targets of the target platform; performing data integration on the original environmental perception data according to the data requirements and the data targets to obtain the environmental perception data that meets the preset conditions.
[0009] Optionally, in an embodiment of this application, the reconstructing the 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data and the target NeRF model includes: based on the target platform, preprocessing the images in the environmental perception data to obtain preprocessed images; sampling the scene points of the preprocessed images in the 3D space, and using the target NeRF model to perform spatio-temporal point interpolation processing on the scene points to obtain the 4D scene corresponding to the environmental perception data.
[0010] Optionally, in an embodiment of this application, the generating and outputting virtual data corresponding to the 4D scene by using the virtual camera includes: generating and rendering a virtual image of the 4D scene in the virtual environment of the virtual camera to obtain a rendered virtual image; adding labels and annotations to the virtual image to obtain the virtual data.
[0011] The second aspect of the present application provides an apparatus for establishing a unified model for vehicle-road-cloud integrated collaborative perception function, which is applied to the cloud. Among them, the apparatus includes: a first establishment module, configured to obtain sensor information of the perception system of the vehicle, and establish a target combination model and a virtual camera of the perception system based on the sensor information and the requirements of the perception system; a processing module, configured to obtain the environmental perception data of the perception system, transmit the environmental perception data to a target platform by using the target combination model, and reconstruct a 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data, and a target NeRF model, so as to generate and output virtual data corresponding to the 4D scene by using the virtual camera; a second establishment module, configured to collect the virtual data to train and fine-tune a target inference model to obtain a final inference model, so as to use the final inference model to infer the perception result of the environmental perception data, and transmit the perception result to the vehicle, so as to complete the establishment of the unified model for vehicle-road-cloud integrated collaborative perception function.
[0012] Optionally, in an embodiment of the present application, the first establishment module includes: a first establishment unit, configured to collect processing module information corresponding to the sensor information, and establish the target combination model according to the processing module information; a second establishment unit, configured to construct a perception scheme according to the sensor information, and establish the virtual camera through the perception scheme.
[0013] Optionally, in an embodiment of the present application, it further includes: an acquisition module, configured to obtain the data requirements and data targets of the target platform before transmitting the environmental perception data to the target platform by using the target combination model; an integration module, configured to perform data integration on the original environmental perception data according to the data requirements and the data targets to obtain the environmental perception data that meets the preset conditions.
[0014] Optionally, in an embodiment of the present application, the processing module includes: a preprocessing unit, configured to preprocess the images in the environmental perception data based on the target platform to obtain preprocessed images; a first processing unit, configured to sample the scene points of the preprocessed images in the 3D space, and perform spatio-temporal point interpolation processing on the scene points by using the target NeRF model to obtain a 4D scene corresponding to the environmental perception data.
[0015] Optionally, in an embodiment of the present application, the processing module includes: a generation unit, configured to generate and render a virtual image of the 4D scene in the virtual environment of the virtual camera to obtain a rendered virtual image; a second processing unit, configured to add labels and annotations to the virtual image to obtain the virtual data.
[0016] A third aspect embodiment of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function as described in the above embodiments.
[0017] A fourth aspect embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the program is executed by a processor, it implements the above method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function.
[0018] A fifth aspect embodiment of the present application provides a computer program product including a computer program, and when the computer program is executed, it is used to implement the above method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function.
[0019] Embodiments of the present application can deploy a model in the cloud based on sensor information of the vehicle's perception system, thereby reconstructing a 4D scene and virtual data corresponding to the environmental perception data in the cloud, training and fine-tuning a target inference model using the virtual data to obtain a final inference model to infer scene information in the environmental perception data, and transmitting the scene information to the vehicle to complete the establishment of a unified model of vehicle-road-cloud integrated collaborative perception function. Thus, the reconstruction of the 4D scene by fusing multi-source data information is realized, meeting the requirements of "multiple vehicle models / multiple platforms / multiple functions". At the same time, the robustness of the NeRF model to data is enhanced. Moreover, the final inference model can infer the perception result of the environmental perception data and send it to the vehicle end, optimizing the inference time while realizing the fusion of vehicle-end and roadside data in the cloud, greatly reducing the computing pressure on the vehicle-end inference device. The real-time transmission of the perception information around the vehicle from the cloud is more conducive to the decision-making of downstream planning and control for autonomous driving. Finally, the establishment of a unified model of vehicle-road-cloud integrated collaborative perception function that can adapt to "multiple vehicle models / multiple platforms / multiple functions" is completed, which can systematically fuse the perception information of different vehicle models, different road conditions, and cloud platforms and perform precise perception on the cloud platform. Thus, it solves the problems that there are still some limitations in the complex task adaptability and performance improvement of in-vehicle perception algorithms in the related art. For example, the position distributions of sensors of different vehicle models and roadside devices are relatively irregular, resulting in a model being difficult to meet the requirements of "multiple vehicle models / multiple platforms / multiple functions", and due to the huge amount of fused data, it is also difficult for the vehicle-end inference device to meet this requirement.
[0020] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0022] Figure 1 It is a flowchart of a method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to an embodiment of the present application;
[0023] Figure 2 It is a schematic framework diagram of the process for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to an embodiment of the present application;
[0024] Figure 3 It is a schematic structural diagram of a device for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to an embodiment of the present application;
[0025] Figure 4 It is a schematic structural diagram of an electronic device according to an embodiment of the present application.
[0026] Reference numerals:
[0027] 10 - Device for establishing a unified model of vehicle-road-cloud integrated collaborative perception function: 100 - First establishment module, 200 - Processing module, and 300 - Second establishment module; 401 - Memory, 402 - Processor, and 403 - Communication interface. Detailed implementation manners
[0028] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0029] The method and device for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the limitations in the adaptability of complex tasks and performance improvement of the vehicle-mounted perception algorithm in the related art mentioned in the above background art, for example, the position distributions of sensors of different vehicle models and roadside devices are relatively irregular, resulting in a model being difficult to adapt to the requirements of "multiple vehicle models / multiple platforms / multiple functions", and due to the huge amount of fused data, it is also difficult for the vehicle-side inference engine to meet this requirement. The present application provides a method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function. In this method, the model can be deployed in the cloud based on the sensor information of the vehicle's perception system, thereby reconstructing the 4D scene and virtual data corresponding to the environmental perception data in the cloud, training and fine-tuning the target inference model with the virtual data to obtain the final inference model to infer the scene information in the environmental perception data, and transmitting the scene information to the vehicle to complete the establishment of the unified model of vehicle-road-cloud integrated collaborative perception function. Thus, the reconstruction of the 4D scene by fusing multi-source data information is realized, meeting the requirements of "multiple vehicle models / multiple platforms / multiple functions", and at the same time, the robustness of the NeRF model to data is enhanced. Moreover, the final inference model can infer the perception results of the environmental perception data and send them to the vehicle side, optimizing the inference time while realizing the fusion of vehicle-side and roadside data in the cloud, greatly reducing the computing pressure on the vehicle-side inference engine. The real-time transmission of the perception information around the vehicle from the cloud is more conducive to the decision-making of downstream planning and control of autonomous driving. Finally, the establishment of a unified model of vehicle-road-cloud integrated collaborative perception function that can adapt to "multiple vehicle models / multiple platforms / multiple functions" is completed, which can systematically fuse the perception information of different vehicle models, different road conditions and cloud platforms, and perform accurate perception on the cloud platform. Thus, the problems in the related art that there are still some limitations in the adaptability of complex tasks and performance improvement of the vehicle-mounted perception algorithm, such as the relatively irregular position distributions of sensors of different vehicle models and roadside devices, resulting in a model being difficult to adapt to the requirements of "multiple vehicle models / multiple platforms / multiple functions", and due to the huge amount of fused data, it is also difficult for the vehicle-side inference engine to meet this requirement, etc., are solved.
[0030] Before explaining the method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to the embodiments of the present application, the computing basic platform involved in the embodiments of the present application will be explained first.
[0031] Among them, the Computing Brain DEvelopment System (CBDES) consists of two parts: the Computing Brain (CBB) and the Graphical ADAS-AD Software Developer (GAASD). The CBB consists of computing platform hardware, a real-time kernel, middleware, and functional software. The functional software is the core of this product, aiming to provide basic algorithm components and frameworks for various intelligent driving systems. Based on the innovative solution of "layered decoupling and cross-domain sharing", with the functional software library as the basis, CBDES cedes the development ability of application algorithms to the host manufacturers, supports the engineers of the host factory to quickly build their own defined intelligent driving systems, and perform function adaptation and parameter tuning. Compared with other products and development models, it has three major advantages of "high efficiency / high quality / generative". The technology involved in the embodiments of this application is mainly used in the vehicle-road-cloud collaborative perception part of the autonomous driving perception model function module in the CBB functional software, aiming to use independent module combinations to achieve multi-vehicle / multi-platform / multi-function, and perform pre-training and fine-tuning.
[0032] Specifically, Figure 1 It is a flowchart of a method for establishing a unified model for vehicle-road-cloud integrated collaborative perception function provided by the embodiments of this application.
[0033] As Figure 1 shown, this method for establishing a unified model for vehicle-road-cloud integrated collaborative perception function is applied to the cloud and includes the following steps:
[0034] In step S101, sensor information of the vehicle's perception system is obtained, and a target combination model and a virtual camera of the perception system are established based on the sensor information and the requirements of the perception system.
[0035] In some embodiments, the vehicle perception system includes a large number of types and quantities of sensors with different functions. Therefore, when establishing a unified model for vehicle-road-cloud integrated collaborative perception in this application, in order to ensure the generalization of the model and be applicable to various vehicles, the sensor information of the vehicle perception system can be obtained first, and a certain combination model and virtual camera can be established based on the sensor information and the functional requirements of the perception system.
[0036] For example, embodiments of the present application can obtain the types and quantities of vehicle sensors. For example, the types of vehicle sensors include, but are not limited to, radar, cameras, ultrasonic sensors, lidar, and other environmental perception devices. The quantity and types of these sensors usually depend on the specific type of vehicle and the required functions. Also, roadside sensors play a key role in traffic management and safety. For example, roadside sensors include, but are not limited to, traffic monitoring cameras, vehicle identification systems, road condition detectors, environmental sensors, etc. Through roadside sensors, traffic flow, road conditions, climate conditions, etc. can be monitored in real time, so as to make real-time adjustments and optimizations to achieve all-round and high-precision environmental perception.
[0037] After obtaining the sensor information, embodiments of the present application can determine multiple independent functional modules according to the requirements of the perception system, that is, the functions that need to be utilized when the perception system processes the information obtained by these sensors, and establish a target combination model by deploying these functional modules in the cloud. Here, the target combination model can be understood as a combination model in the vehicle networking and intelligent transportation systems that combines vehicle-related functional modules with the cloud platform to realize real-time collection, processing, analysis, and remote management and control of data.
[0038] The cloud deployment model is a key concept in modern information technology, which provides enterprises and developers with a flexible, scalable, and efficient way to run and manage applications, services, and data. The advantage of the cloud deployment model lies in its ability to utilize cloud computing resources to provide elasticity, security, and reliability. First of all, the cloud deployment model eliminates the need for traditional physical servers and data centers, thus reducing the maintenance cost and complexity. Enterprises can dynamically adjust computing, storage, and network resources according to actual needs, avoiding over-investment or resource shortages. In addition, the development of cloud computing and distributed computing environments has also brought a significant improvement in the model operation speed.
[0039] Furthermore, embodiments of the present application can also determine the relevant information of the virtual camera that needs to be utilized in the subsequent processing process according to these sensor information, so as to establish a certain virtual camera.
[0040] Embodiments of the present application can collect the combination of sensor data and corresponding independent functional modules, deploy them in the cloud, establish a certain combination model, and determine various information of the virtual camera, which completes the preparatory stage of vehicle-road-cloud integration. This process fully reflects the satisfaction of the model for the requirements of "multiple vehicle types / multiple platforms / multiple functions".
[0041] Optionally, in an embodiment of the present application, establishing a target combination model and a virtual camera based on sensor information and the requirements of the perception system includes: collecting processing module information corresponding to the sensor information to establish a target combination model according to the processing module information; constructing a perception scheme based on the sensor information to establish a virtual camera through the perception scheme.
[0042] Based on the relevant descriptions of other embodiments, it can be understood that the embodiments of the present application can establish a target combination model and a virtual camera according to the sensor information and the requirements of the perception system.
[0043] In the actual execution process, the present application mainly but not limited to first collecting the processing module information corresponding to the sensor information, that is, the independent functional modules required when processing the information acquired by the sensor, and then combining these independent functional modules and uploading them to the cloud to establish a target combination model according to the combined independent functional modules. And, constructing a perception scheme based on these sensor information to establish a virtual camera.
[0044] When determining the independent functional modules, the embodiments of the present application can select an appropriate combination of independent functional modules according to the requirements and objectives of the perception system to achieve specific functions and tasks. These modules include but are not limited to a data preprocessing module, a feature extraction module, a classification module, an identification module, etc. Then, considering the complexity and variability of the actual application, the requirements of different scenarios and environments, as well as the scalability, flexibility, and adaptability of the perception system, these independent modules are combined and optimized, thereby realizing the efficient operation and performance improvement of the perception system.
[0045] It should be noted that the specific selection and combination of independent functional modules can be determined and adjusted by those skilled in the art according to the actual situation. Only exemplary descriptions are given in the embodiments of the present application without specific limitations.
[0046] And, the embodiments of the present application can establish a certain perception scheme based on the advantages of vehicle-mounted sensors and roadside sensors, such as a wide range and high precision, so as to improve the overall efficiency and safety of the traffic system. After establishing the perception scheme, the embodiments of the present application can stipulate the number of virtual cameras, parameters, etc. according to the requirements of "multiple vehicle types / multiple platforms / multiple functions", thereby establishing a virtual camera.
[0047] Step S102, obtaining the environmental perception data of the perception system, using the target combination model to transmit the environmental perception data to the target platform, and reconstructing the 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data, and the target NeRF model, so as to generate and output virtual data corresponding to the 4D scene using the virtual camera.
[0048] As a possible implementation, when constructing a unified model for the collaborative perception function of vehicle-road-cloud integration in the embodiments of the present application, the important functions of the perception system are mainly but not limited to being realized by reconstructing a three-dimensional scene into a 4D scene through 4D reconstruction. When reconstructing the three-dimensional scene, the embodiments of the present application mainly first obtain environmental perception data acquired by sensors such as vehicle-mounted sensors and roadside sensors, then transmit this environmental perception data to a certain computing platform or processing platform through a combined model deployed in the cloud, and then reconstruct the 4D scene in combination with this platform and the target NeRF model, and finally generate and output virtual data corresponding to the 4D scene using a virtual camera.
[0049] Next, a further explanation of this process will be given.
[0050] Optionally, in an embodiment of the present application, before transmitting the environmental perception data to the target platform using the target combined model, it further includes: obtaining the data requirements and data targets of the target platform; performing data integration on the original environmental perception data according to the data requirements and data targets to obtain environmental perception data that meets preset conditions.
[0051] In some embodiments, before transmitting the environmental perception data to a certain target platform for processing using the target combined model, it is necessary to first obtain the data requirements and data targets of the target platform. Here, the target platform can be understood as a platform for performing various processes on the environmental perception data. For example, the CBB platform; the data requirements can be understood as certain requirements that the data needs to meet when the target platform processes the data. For example, the data has security, integrity, etc.; the data target can be understood as certain targets that the data needs to meet after the data processing is completed. For example, the data has confidentiality, accuracy, etc.
[0052] Then, perform data integration on the original environmental perception data according to the data requirements and data targets to obtain environmental perception data that meets preset conditions. Here, the preset conditions can be understood as certain conditions that the original environmental perception data needs to meet before being input into the target platform. For example, the format is correct and within the valid data range, etc.
[0053] For example, the present application can install a CBB (Common Building Block) platform in the cloud as the target platform. It can use cloud computing technology to deploy the CBB component library, management platform and related tools on the cloud server, and provide a platform for R & D personnel to remotely access and use services through the Internet or the enterprise internal network. And this platform can support the collaborative work of multiple R & D teams, realizing the efficient reuse and rapid iteration of CBB components.
[0054] Next, the embodiments of the present application can integrate the original environmental perception data from different sources and formats into a unified data structure or pattern, and use data fusion technologies and methods (such as data warehouses, ETL tools, etc.) to ensure the consistency, integrity, and accuracy of the data. Then, the processed environmental perception data is transmitted to the CBB platform.
[0055] Finally, according to the data requirements of the CBB platform, for example, the data is further processed and optimized, such as feature selection, dimensionality reduction, data augmentation, etc., and considering the requirements of data security and privacy protection, data encryption, desensitization, or anonymization processing is performed; then, based on the data target, such as ensuring the secure, private, and efficient transmission of the data to the CBB platform while improving the data quality, the appropriate data transmission protocol and method are selected to transmit the packaged data to the CBB platform securely and efficiently.
[0056] Optionally, in an embodiment of the present application, reconstructing a 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data, and the target NeRF model includes: preprocessing the images in the environmental perception data based on the target platform to obtain the preprocessed images; sampling the scene points of the preprocessed images in the 3D space, and using the target NeRF model to perform spatio-temporal point interpolation processing on the scene points to obtain the 4D scene corresponding to the environmental perception data.
[0057] In some embodiments, when reconstructing a 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data, and the target NeRF model, the present application can, but is not limited to, using the target platform to preprocess the images in the environmental perception data, then sampling the scene points of the preprocessed images in the 3D space, and finally using the target NeRF model to perform spatio-temporal point interpolation processing on the scene points to obtain the 4D scene corresponding to the environmental perception data. Here, the target NeRF model can be understood as various NeRF models for scene reconstruction.
[0058] For example, the embodiments of the present application can first use the CBB platform to preprocess the images, such as cropping, scaling, and color correction, which helps to improve the training effect of the NeRF model. Then, uniformly sample the scene points in the images in the 3D space, and considering the discrete step size in the time dimension, use the trained NeRF model to interpolate the spatio-temporal points to reconstruct the dense representation of the 4D scene.
[0059] Additionally, the embodiments of the present application can perform qualitative and quantitative evaluations on the reconstructed 4D scene to evaluate the performance and reconstruction quality of the NeRF model, and optimize and tune the NeRF model according to the evaluation results to improve the accuracy, stability, and efficiency of the 4D scene reconstruction.
[0060] The embodiments of this application can fuse multi-source data information for the reconstruction of 4D scenarios, meet the requirements of "multiple vehicle models / multiple platforms / multiple functions", and enhance the robustness of the NeRF model to data.
[0061] Optionally, in an embodiment of this application, virtual data corresponding to a 4D scenario is generated and output using a virtual camera, including: generating and rendering a virtual image of the 4D scenario in the virtual environment of the virtual camera to obtain the rendered virtual image; adding annotations and comments to the virtual image to obtain virtual data.
[0062] In some other embodiments, when generating and outputting virtual data corresponding to a 4D scenario using a virtual camera, this application can generate and render a virtual image corresponding to the 4D scenario in the virtual environment of the virtual camera, and add corresponding annotations and comments to the rendered virtual image to obtain and output virtual data of a fixed quantity and category.
[0063] For example, after establishing a virtual camera based on sensor information, that is, the number and its parameters of the virtual camera have been set, this application can first use the graphics rendering engine (such as Unity, Unreal Engine, etc.) in the virtual camera to generate and render a virtual image corresponding to the 4D scenario in the virtual environment to obtain the rendered virtual image.
[0064] Then, the embodiments of this application can annotate and comment on the generated virtual data to identify and classify different categories and attributes. For example, static objects - traffic signs, dynamic objects - vehicles, etc. Finally, the rendered and annotated virtual data is exported in common image or video formats (such as PNG, JPEG, MP4, etc.).
[0065] Thus, the embodiments of this application can obtain the environmental perception data of all sensors and perform 4D scenario reconstruction, and finally output the virtual camera data corresponding to the 4D scenario. Through these processes and the combined independent module parts, the model meets the requirements of "multiple vehicle models / multiple platforms / multiple functions".
[0066] Step S103, collect virtual data to train and fine-tune the target inference model to obtain the final inference model, use the final inference model to infer the perception result of the environmental perception data, and transmit the perception result to the vehicle to complete the establishment of the unified model for vehicle-road-cloud integrated collaborative perception function.
[0067] As a possible implementation method, after obtaining the virtual data output by the virtual camera, the embodiments of the present application can use the virtual data to train and fine-tune the target inference model, so as to use the final inference model to infer the perception result of the environment perception data during the actual application process, and then transmit the perception result back to the vehicle, thereby completing the establishment of the unified model for the vehicle-road-cloud integrated collaborative perception function. Herein, the target inference model can be understood as an inference model that has been pre-trained.
[0068] Based on the virtual data and the pre-trained inference model, the embodiments of the present application can further use the virtual data to train the pre-trained inference model and perform fine-tuning, thereby optimizing the performance of the inference model to obtain the final inference model. During the fine-tuning process, the embodiments of the present application can, but are not limited to, freeze the first few layers or all layers of the pre-trained model, and only update specific layers or newly added layers to retain the features learned by the pre-trained model for model fine-tuning, and this processing can accelerate the fine-tuning process and implement different deep learning strategies as needed.
[0069] Through fine-tuning, the embodiments of the present application can effectively utilize limited data resources to help the model adapt to new data distributions and characteristics, improve the task adaptability of the final inference model, and further improve the performance of the final inference model on specific tasks.
[0070] After using the final inference model to infer the perception result of the environment perception data, it is necessary to transmit the perception result to the vehicle side, thereby completing the establishment of the unified model for the vehicle-road-cloud integrated collaborative perception function.
[0071] In summary, cloud platform computing can integrate all data and obtain great benefits in model inference, which meets the requirement of optimizing the inference time of the model and enhances the vehicle's environmental perception ability.
[0072] Additionally, the embodiments of the present application can also use virtual data to train the combined independent modules, so that the combined model can quickly conform to the scenario application.
[0073] The embodiments of the present application can construct a final inference model that can protect privacy while being able to respond to tasks in real time, and the final inference model can reduce bandwidth requirements, improve the offline processing ability, thereby reducing the inference time. Finally, transmitting the perception result to the vehicle side can reduce the computing pressure on the vehicle-side inference device and enable the vehicle side to obtain surrounding perception information in real time, which helps the vehicle side with planning and control.
[0074] Next, a specific embodiment is used to elaborate in detail on the method for establishing the unified model for the vehicle-road-cloud integrated collaborative perception function in the embodiments of the present application.
[0075] Figure 2This is a framework schematic diagram of the establishment process of the unified model for vehicle-road-cloud integrated collaborative perception function in an embodiment of this application. As Figure 2 shown:
[0076] (1) Investigate the types and quantities of vehicle-mounted sensors;
[0077] (2) Determine the perception scheme based on vehicle-mounted and roadside sensors;
[0078] (3) Select the independent module combination based on the perception scheme;
[0079] (4) Deploy the model in the cloud;
[0080] (5) Obtain the environmental perception data of all sensors;
[0081] (6) Package all the environmental perception data and input it into the CBB platform;
[0082] (7) Implement 4D scene reconstruction of the environmental perception data based on the NeRF model;
[0083] (8) Output virtual data corresponding to the 4D scene with a fixed quantity and category based on the virtual camera;
[0084] (9) Collect the virtual data to train the combined model to enhance the scene application effect of the combined model, and at the same time input the virtual data into the pre-trained inference model to obtain the perception result;
[0085] (10) Output the perception result to the vehicle end;
[0086] (11) Use the above task process to fine-tune the pre-trained inference model to obtain the final inference model, and use the final inference model after fine-tuning for inference in the next task to obtain a more accurate perception result.
[0087] According to the method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function proposed in the embodiments of the present application, the model can be deployed in the cloud based on the sensor information of the perception system of the vehicle, thereby reconstructing the 4D scene and virtual data corresponding to the environmental perception data in the cloud, training and fine-tuning the target inference model with the virtual data to obtain the final inference model to infer the scene information in the environmental perception data, and transmitting the scene information to the vehicle to complete the establishment of the unified model of the vehicle-road-cloud integrated collaborative perception function. Thus, the reconstruction of the 4D scene by fusing multi-source data information is realized, meeting the requirements of "multiple vehicle models / multiple platforms / multiple functions". At the same time, the robustness of the NeRF model to data is enhanced. Moreover, the final inference model can infer the perception results of the environmental perception data and send them to the vehicle side, optimizing the inference time while realizing the fusion of vehicle-side and roadside data in the cloud, greatly reducing the computing pressure on the vehicle-side inference device. The real-time perception information transmitted by the cloud to the surrounding of the vehicle is more conducive to the decision-making of downstream planning and control for autonomous driving. Finally, the establishment of a unified model of the vehicle-road-cloud integrated collaborative perception function that can adapt to "multiple vehicle models / multiple platforms / multiple functions" is completed, which can systematically fuse the perception information of different vehicle models, different road conditions, and cloud platforms and perform accurate perception on the cloud platform. Thus, it solves the problems that there are still some limitations in the complex task adaptability and performance improvement of the in-vehicle perception algorithm in the related technology. For example, the position distributions of sensors of different vehicle models and roadside devices are relatively irregular, resulting in a model being difficult to adapt to the requirements of "multiple vehicle models / multiple platforms / multiple functions", and due to the huge amount of fused data, it is also difficult for the vehicle-side inference device to meet this requirement, etc.
[0088] Next, a device for establishing a unified model of vehicle-road-cloud integrated collaborative perception function proposed in the embodiments of the present application will be described with reference to the accompanying drawings.
[0089] Figure 3 It is a schematic structural diagram of a device for establishing a unified model of vehicle-road-cloud integrated collaborative perception function according to an embodiment of the present application.
[0090] As Figure 3 shown, the device 10 for establishing a unified model of vehicle-road-cloud integrated collaborative perception function is applied to the cloud and includes: a first establishment module 100, a processing module 200, and a second establishment module 300.
[0091] Among them, the first establishment module 100 is used to obtain the sensor information of the perception system of the vehicle and establish a target combination model and a virtual camera of the perception system based on the sensor information and the requirements of the perception system.
[0092] The processing module 200 is configured to obtain the environmental perception data of the perception system, transmit the environmental perception data to the target platform by using the target combination model, and reconstruct the 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data, and the target NeRF model, so as to generate and output virtual data corresponding to the 4D scene by using a virtual camera.
[0093] The second establishment module 300 is configured to collect virtual data to train and fine-tune the target inference model to obtain a final inference model, and use the final inference model to infer the perception result of the environmental perception data, and transmit the perception result to the vehicle to complete the establishment of the unified model for vehicle-road-cloud integrated collaborative perception function.
[0094] Optionally, in an embodiment of the present application, the first establishment module 100 includes: a first establishment unit and a second establishment unit.
[0095] The first establishment unit is configured to collect the processing module information corresponding to the sensor information, and establish a target combination model according to the processing module information.
[0096] The second establishment unit is configured to construct a perception scheme according to the sensor information, and establish a virtual camera through the perception scheme.
[0097] Optionally, in an embodiment of the present application, it further includes: an acquisition module and an integration module.
[0098] The acquisition module is configured to obtain the data requirements and data targets of the target platform before transmitting the environmental perception data to the target platform by using the target combination model.
[0099] The integration module is configured to perform data integration on the original environmental perception data according to the data requirements and data targets to obtain environmental perception data that meets the preset conditions.
[0100] Optionally, in an embodiment of the present application, the processing module 200 includes: a preprocessing unit and a first processing unit.
[0101] The preprocessing unit is configured to preprocess the images in the environmental perception data based on the target platform to obtain preprocessed images.
[0102] The first processing unit is configured to sample the scene points of the preprocessed images in the 3D space, and perform spatio-temporal point interpolation processing on the scene points by using the target NeRF model to obtain a 4D scene corresponding to the environmental perception data.
[0103] Optionally, in an embodiment of the present application, the processing module 200 includes: a generation unit and a second processing unit.
[0104] Among them, a generation unit is configured to generate and render a virtual image of a 4D scene in a virtual environment of a virtual camera to obtain a rendered virtual image.
[0105] A second processing unit is configured to add annotations and comments to the virtual image to obtain virtual data.
[0106] It should be noted that the foregoing explanation of the method embodiments for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function also applies to the device for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function in this embodiment, and will not be repeated here.
[0107] According to the device for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function provided by an embodiment of the present application, a model can be deployed in the cloud based on sensor information of the perception system of a vehicle, thereby reconstructing a 4D scene and virtual data corresponding to environmental perception data in the cloud, training and fine-tuning a target inference model using the virtual data to obtain a final inference model to infer scene information in the environmental perception data, and transmitting the scene information to the vehicle to complete the establishment of the unified model of the vehicle-road-cloud integrated collaborative perception function. Thus, the reconstruction of the 4D scene by fusing multi-source data information is realized, meeting the requirements of "multiple vehicle models / multiple platforms / multiple functions". At the same time, the robustness of the NeRF model to data is enhanced. Moreover, the final inference model can infer the perception result of the environmental perception data and send it to the vehicle end, optimizing the inference time while realizing the fusion of vehicle-end and roadside data in the cloud, greatly reducing the computing pressure on the vehicle-end inference device. And the perception information transmitted by the cloud to the surrounding of the vehicle in real time is more conducive to the decision-making of downstream planning and control for autonomous driving. Finally, the establishment of a unified model of the vehicle-road-cloud integrated collaborative perception function that can adapt to "multiple vehicle models / multiple platforms / multiple functions" is completed, which can systematically fuse the perception information of different vehicle models, different road conditions and cloud platforms, and perform accurate perception on the cloud platform. Thus, it solves the problems that there are still some limitations in the complex task adaptability and performance improvement of in-vehicle perception algorithms in the related art. For example, the position distributions of sensors of different vehicle models and roadside devices are relatively irregular, resulting in a model being difficult to adapt to the requirements of "multiple vehicle models / multiple platforms / multiple functions", and due to the huge amount of fused data, it is also difficult for the vehicle-end inference device to meet this requirement.
[0108] Figure 4 The structural schematic diagram of the electronic device provided by an embodiment of the present application. The electronic device may include:
[0109] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.
[0110] When the processor 402 executes the program, it implements the method for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function provided in the foregoing embodiment.
[0111] Furthermore, the electronic device further includes:
[0112] A communication interface 403 for communication between the memory 401 and the processor 402.
[0113] A memory 401 for storing computer programs that can run on the processor 402.
[0114] The memory 401 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0115] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0116] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.
[0117] The processor 402 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0118] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function is implemented.
[0119] The embodiment of the present application also provides a computer program product, including a computer program that can run computer instructions, and when the computer instructions are executed by a processor, the unified model establishment method for vehicle-road-cloud integrated collaborative perception function provided by the embodiment of the present application is implemented.
[0120] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0121] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0122] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of the code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiment of the present application includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in the order, including in a substantially simultaneous manner or in the reverse order according to the functions involved, which should be understood by those skilled in the art of the embodiments of the present application.
[0123] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0124] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0125] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0126] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0127] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for establishing a unified model of vehicle-road-cloud integrated collaborative perception function, characterized in that: Applied to the cloud, wherein the method comprises the following steps: Acquire sensor information of a perception system of the vehicle, and establish a target combination model and a virtual camera of the perception system based on the sensor information and requirements of the perception system; Acquire environmental perception data of the perception system, transmit the environmental perception data to a target platform using the target combination model, reconstruct a 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data and a target NeRF model, and generate and output virtual data corresponding to the 4D scene using the virtual camera; The virtual data is collected to train and fine-tune the target reasoning model to obtain a final reasoning model, and the final reasoning model is used to infer the perception results of the environmental perception data, and the perception results are transmitted to the vehicle to complete the establishment of a unified model of vehicle-road-cloud integrated collaborative perception function.
2. The method according to claim 1, characterized in that The establishing of the target combination model and the virtual camera of the perception system based on the sensor information and the requirements of the perception system includes: Collecting processing module information corresponding to the sensor information to establish the target combination model according to the processing module information; A perception scheme is constructed according to the sensor information to establish the virtual camera through the perception scheme.
3. The method according to claim 1, characterized in that Before transmitting the environment perception data to the target platform using the target combination model, the method further includes: Obtaining data requirements and data targets of the target platform; The original environmental perception data is integrated according to the data requirements and the data targets to obtain the environmental perception data that meets preset conditions.
4. The method according to claim 1, characterized in that: The reconstructing the 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data and the target NeRF model includes: Based on the target platform, preprocessing the image in the environmental perception data to obtain a preprocessed image; The scene points of the preprocessed image in the 3D space are sampled, and the scene points are subjected to space-time point interpolation processing using the target NeRF model to obtain a 4D scene corresponding to the environmental perception data.
5. The method according to claim 1, characterized in that The step of generating and outputting virtual data corresponding to the 4D scene using the virtual camera includes: Generating and rendering a virtual image of the 4D scene in a virtual environment of the virtual camera to obtain a rendered virtual image; Adding annotations and comments to the virtual image to obtain the virtual data.
6. A unified model building device for vehicle-road-cloud integrated collaborative perception function, characterized in that: Applied to the cloud, wherein the device comprises: A first establishing module is used to obtain sensor information of a perception system of a vehicle, and to establish a target combination model and a virtual camera of the perception system based on the sensor information and requirements of the perception system; a processing module, for acquiring environmental perception data of the perception system, transmitting the environmental perception data to a target platform using the target combination model, and reconstructing a 4D scene corresponding to the environmental perception data based on the target platform, the environmental perception data and the target NeRF model, so as to generate and output virtual data corresponding to the 4D scene using the virtual camera; The second establishment module is used to collect the virtual data training and fine-tune the target reasoning model to obtain a final reasoning model, and use the final reasoning model to infer the perception results of the environmental perception data, and transmit the perception results to the vehicle to complete the establishment of a unified model of vehicle-road-cloud integrated collaborative perception function.
7. The device according to claim 6, characterized in that The first establishing module comprises: A first establishing unit, configured to collect processing module information corresponding to the sensor information, so as to establish the target combination model according to the processing module information; The second establishing unit is used to construct a perception scheme according to the sensor information to establish the virtual camera through the perception scheme.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement a method for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function as described in any one of claims 1-5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, it is used to implement the method for establishing a unified model of the vehicle-road-cloud integrated collaborative perception function as described in any one of claims 1-5.
Citation Information
Cited By
Target identification method and device and vehicle and road cloud integrated test system
CN122410508A