Intelligent network connection environment-oriented vehicle-road cloud real-time collaborative perception method and device

By acquiring vehicle-road point cloud data in an intelligent connected environment and fusing it at the edge and cloud, the perception model is dynamically optimized, which solves the problems of insufficient communication latency modeling and model adaptability in vehicle-road cooperative perception algorithms, and improves the accuracy of real-time perception.

CN119743498BActive Publication Date: 2025-10-24TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411560887.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-10-24
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Existing vehicle-road cooperative perception algorithms fail to effectively model the communication latency of perceptual information transmission to the cloud, resulting in a situation where computational errors and latency mutually constrain each other in the cloud control system of intelligent connected vehicles, and the system cannot be adjusted according to the real-time dynamic environment to adapt to the needs of different environmental complexities.

Method used

By acquiring vehicle-side and roadside cloud data within the target traffic scenario, generating perception results using vehicle-side and roadside perception models, and fusing them in the edge cloud, calculating transmission latency and fusion duration, and dynamically optimizing the perception model based on real-time perception performance indicators, real-time collaborative perception between vehicles, roads, and the cloud is achieved.

Benefits of technology

It improves the real-time accuracy of the target-level collaborative perception algorithm, solves the problems of insufficient communication latency modeling and model adaptability, and realizes dynamic adjustment to different environmental complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743498B_ABST
    Figure CN119743498B_ABST
Patent Text Reader

Abstract

The application relates to a vehicle-road cloud real-time collaborative perception method and device for an intelligent network connection environment, wherein the method comprises the following steps: acquiring multiple frames of vehicle-end and road-end point cloud data of vehicle-end and road-end information nodes, inputting the vehicle-end and road-end perception models to generate vehicle-end and road-end perception results, and fusing the perception results through an edge cloud; acquiring the vehicle cloud and road cloud transmission delays between the vehicle-end and road-end and the edge cloud to calculate the total length of the fusion output, and combining the perception output lag parameters to calculate a real-time perception performance index; based on the real-time perception performance index and a real-time dynamic algorithm optimization strategy, target optimization vehicle-end and road-end perception models are obtained to generate vehicle-road cloud real-time collaborative perception results. Therefore, the problems that the prior art cannot model the communication delay of the transmission of perception information to the cloud in the intelligent network connection automobile cloud control system from the algorithm perspective, the collaborative perception model has mutual restrictions of calculation errors and delays, and different models of different scales need to be used for different environment complexities are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle-road cloud cooperative perception, and particularly relates to a vehicle-road cloud real-time cooperative perception method and device for an intelligent networked environment. BACKGROUND

[0002] An intelligent networked vehicle realizes intelligent information sharing between vehicles, roads, and clouds through on-vehicle sensors, controllers, actuators, and the like in combination with modern communication and network technologies. An intelligent networked vehicle cloud control system is a closed-loop control process in which a real-world intelligent vehicle and a driving environment are mapped and digitally twinned in the cloud through perception and positioning communication technologies, safe, energy-saving, and efficient functions are applied in the cloud, and instructions are issued. In this process, the physical world and the information world are coupled, and the physical side and the information side are coupled. In this system, perception needs to have low latency in communication and calculation and high accuracy in mapping the physical world to the information world.

[0003] However, existing vehicle-road cooperative perception algorithms mainly focus on the offline performance of the algorithm and ignore the real-time problem of the algorithm in actual application, resulting in the following problems in the digital twinning of the intelligent networked vehicle cloud control system and other real-time applications:

[0004] 1. The communication latency of the perception information transmitted to the cloud in the intelligent networked vehicle cloud control system is not modeled from the algorithmic perspective.

[0005] The current cooperative perception algorithm design only gives the information interaction relationship and algorithm processing process for different perception nodes, and less considers the cooperative perception algorithm with cloud participation. Moreover, the influence of communication latency on the processing latency of specific applications in the vehicle-road-cloud integrated system is not considered, and no explicit solution is given for the instability of cloud communication.

[0006] 2. The cooperative perception model has a mutual restriction between calculation error and latency error, and different model scales need to be dynamically optimized for different environmental complexities.

[0007] The existing cooperative perception method based on a target-level algorithm has the advantage of good real-time performance and is more suitable for intelligent vehicle cyber-physical systems, but its perception accuracy is generally poor, and the algorithm cannot be adjusted according to real-time dynamic environments after deployment, and the performance is poor.

[0008] In summary, the existing technology cannot model the communication latency of the perception information transmitted to the cloud in the intelligent networked vehicle cloud control system from the algorithmic perspective, and the cooperative perception model has a mutual restriction between calculation error and latency error. Different model scales are needed for different environmental complexities, which needs to be solved urgently. SUMMARY

[0009] The application provides a vehicle-road cloud real-time collaborative perception method and device for an intelligent network environment, to solve the problem that the prior art cannot model the communication delay of perception information transmission to the cloud in the intelligent network vehicle cloud control system from the algorithm perspective, and the collaborative perception model has mutual constraints such as calculation error and delay, and different models of different scales need to be used for different environmental complexities.

[0010] The first aspect embodiment of the application provides a vehicle-road cloud real-time collaborative perception method for an intelligent network environment, including the following steps: acquiring multiple frames of vehicle end point cloud data and multiple frames of road end point cloud data corresponding to at least one target vehicle end information node and at least one road end information node in a target traffic scene respectively, and inputting the multiple frames of vehicle end point cloud data and the multiple frames of road end point cloud data into a preset vehicle end perception model and a road end perception model respectively to generate vehicle end perception results and road end perception results; fusing the vehicle end perception results and the road end perception results through a target edge cloud to obtain a collaborative perception result and a fusion time length, and acquiring a vehicle cloud transmission delay and a road cloud transmission delay between the at least one target vehicle end information node and the at least one road end information node and the target edge cloud respectively, to calculate a total fusion output time length based on the vehicle cloud transmission delay, the road cloud transmission delay and the fusion time length, and calculate a real-time perception performance index corresponding to the collaborative perception result according to the total fusion output time length and a preset perception output lag parameter; performing a preset dynamic optimization operation on the vehicle end perception model and the road end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle end perception model and a target optimized road end perception model, and generating a vehicle-road cloud real-time collaborative perception result according to the target optimized vehicle end perception model and the target optimized road end perception model.

[0011] Optionally, in one embodiment of the application, the acquiring multiple frames of vehicle end point cloud data and multiple frames of road end point cloud data corresponding to at least one target vehicle end information node and at least one road end information node in a target traffic scene respectively, and inputting the multiple frames of vehicle end point cloud data and the multiple frames of road end point cloud data into a preset vehicle end perception model and a road end perception model to generate vehicle end perception results and road end perception results includes: acquiring multiple frames of vehicle end point cloud data and multiple frames of road end point cloud data through the at least one target vehicle end information node and the at least one road end information node respectively, and detecting and tracking the multiple frames of vehicle end point cloud data using the vehicle end perception model to generate the vehicle end perception results; and performing road end target detection and feature extraction operations on the multiple frames of road end point cloud data using the road end perception model to generate the road end perception results.

[0012] Optionally, in an embodiment of the present application, the target edge cloud fuses the vehicle-end perception result and the road-end perception result to obtain a cooperative perception result and a fusion time length, comprising: the target edge cloud performs coordinate conversion processing on the vehicle-end perception result and the road-end perception result, and performs target-level information fusion detection operation on the vehicle-end perception result and the road-end perception result after the coordinate conversion processing, to generate the cooperative perception result and the fusion time length.

[0013] Optionally, in an embodiment of the present application, the vehicle-cloud transmission time delay and the road-cloud transmission time delay between the at least one target vehicle-end information node, the at least one road-end information node and the target edge cloud are respectively acquired, the fusion output total time length is calculated based on the vehicle-cloud transmission time delay, the road-cloud transmission time delay and the fusion time length, and the real-time perception performance index corresponding to the cooperative perception result is calculated according to the fusion output total time length and a preset perception output lag parameter, comprising: determining a road-side model perception time length corresponding to the vehicle-end perception model and a vehicle-side model perception time length corresponding to the road-end perception model; calculating the fusion output total time length according to the road-side model perception time length, the vehicle-side model perception time length, the vehicle-cloud transmission time delay, the road-cloud transmission time delay and the fusion time length; comparing the fusion output total time length with the perception output lag parameter; when the fusion output total time length is less than the perception output lag parameter, calculating a first perception result performance corresponding to each frame of vehicle-end point cloud data in the multi-frame vehicle-end point cloud data; when the fusion output total time length is greater than or equal to the perception output lag parameter, calculating a second perception result performance corresponding to each frame of vehicle-end point cloud data in the multi-frame vehicle-end point cloud data; acquiring a total frame number of the multi-frame vehicle-end point cloud data, and calculating a real-time perception performance total sum corresponding to the multi-frame vehicle-end point cloud data according to the first perception result performance and the second perception result performance, and calculating the real-time perception performance index based on the real-time perception performance total sum and the total frame number.

[0014] Optionally, in an embodiment of the present application, the preset dynamic optimization operation is performed on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, comprising: performing feature engineering environment feature modeling on a target traffic scene corresponding to each frame of cooperative perception result in a plurality of frames of cooperative perception results to obtain an environment feature vector corresponding to the target traffic scene; constructing a plurality of groups of model parameter sets of different scales for the vehicle-end perception model and the road-end perception model respectively, and determining an AP index corresponding to each group of model parameter sets of different scales in the plurality of groups of model parameter sets of different scales; inputting the vehicle-cloud transmission delay, the road-cloud transmission delay, the environment feature vector and the AP index corresponding to a first frame of cooperative perception result in the plurality of frames of cooperative perception results into a pre-trained regression prediction model to generate target model parameters in the plurality of groups of model parameter sets of different scales that meet a preset real-time perception performance index requirement; after a preset time interval, updating the vehicle-cloud transmission delay, the road-cloud transmission delay, the environment feature vector and the AP index according to the vehicle-end perception model and the road-end perception model corresponding to the target model parameters to obtain new vehicle-cloud transmission delay, new road-cloud transmission delay, new environment feature vector and new AP index corresponding to a second frame of cooperative perception result, and inputting the new vehicle-cloud transmission delay, the new road-cloud transmission delay, the new environment feature vector, the new AP index and the target model parameters into the regression prediction model to generate new target model parameters in the plurality of groups of model parameter sets of different scales that meet the preset real-time perception performance index requirement; iteratively performing the updating operation of the vehicle-cloud transmission delay, the road-cloud transmission delay, the environment feature vector, the AP index and the target model parameters until the target model parameters corresponding to each frame of cooperative perception result are generated, so as to obtain the target optimized vehicle-end perception model and the target optimized road-end perception model corresponding to each frame of cooperative perception result based on the target model parameters corresponding to each frame of cooperative perception result.

[0015] The second aspect embodiment of the application provides a vehicle-road cloud real-time collaborative perception device for an intelligent network environment, comprising: a perception module, configured to acquire multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene respectively, and input the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into a preset vehicle-end perception model and a road-end perception model respectively to generate a vehicle-end perception result and a road-end perception result; an index construction module, configured to fuse the vehicle-end perception result and the road-end perception result through a target edge cloud to obtain a collaborative perception result and a fusion time length, acquire a vehicle-cloud transmission time delay and a road-cloud transmission time delay between the at least one target vehicle-end information node and the at least one road-end information node and the target edge cloud respectively, calculate a total fusion output time length based on the vehicle-cloud transmission time delay, the road-cloud transmission time delay and the fusion time length, and calculate a real-time perception performance index corresponding to the collaborative perception result according to the total fusion output time length and a preset perception output lag parameter; and a dynamic optimization module, configured to perform a preset dynamic optimization operation on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and generate a vehicle-road cloud real-time collaborative perception result according to the target optimized vehicle-end perception model and the target optimized road-end perception model.

[0016] Optionally, in an embodiment of the application, the perception module comprises: an acquisition unit, configured to acquire multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data through the at least one target vehicle-end information node and the at least one road-end information node respectively, and detect and track the multi-frame vehicle-end point cloud data by using the vehicle-end perception model to generate the vehicle-end perception result; and a feature extraction unit, configured to perform road-end target detection and feature extraction operation on the multi-frame road-end point cloud data by using the road-end perception model to generate the road-end perception result.

[0017] Optionally, in an embodiment of the application, the index construction module comprises: a fusion unit, configured to perform coordinate conversion processing on the vehicle-end perception result and the road-end perception result through the target edge cloud, and perform target-level information fusion detection operation on the vehicle-end perception result and the road-end perception result after the coordinate conversion processing to generate the collaborative perception result and the fusion time length.

[0018] Optionally, in an embodiment of the present application, the index construction module further comprises: a first determination unit configured to determine a road-side model perception duration corresponding to the vehicle-end perception model and a vehicle-side model perception duration corresponding to the road-end perception model; a first calculation unit configured to calculate the total fusion output duration according to the road-side model perception duration, the vehicle-side model perception duration, the vehicle-cloud transmission delay, the road-cloud transmission delay, and the fusion duration; a comparison unit configured to compare the total fusion output duration with the perception output lag parameter; a second calculation unit configured to calculate a first perception result performance corresponding to each frame of vehicle-end point cloud data in the multi-frame vehicle-end point cloud data when the total fusion output duration is less than the perception output lag parameter; a third calculation unit configured to calculate a second perception result performance corresponding to each frame of vehicle-end point cloud data in the multi-frame vehicle-end point cloud data when the total fusion output duration is greater than or equal to the perception output lag parameter; and a fourth calculation unit configured to obtain a total number of frames of the multi-frame vehicle-end point cloud data, calculate a total real-time perception performance corresponding to the multi-frame vehicle-end point cloud data according to the first perception result performance and the second perception result performance, and calculate the real-time perception performance index based on the total real-time perception performance and the total number of frames.

[0019] Optionally, in an embodiment of the present application, the dynamic optimization module comprises: a modeling unit configured to model a target traffic scene corresponding to each frame of the multi-frame cooperative perception result to obtain an environment feature vector corresponding to the target traffic scene; a second determination unit configured to construct a plurality of sets of model parameters of different scales for the vehicle-side perception model and the road-side perception model respectively, and determine an AP index corresponding to each set of model parameters of different scales in the plurality of sets of model parameters of different scales; a first generation unit configured to input the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, and the AP index corresponding to a first frame of the multi-frame cooperative perception result into a pre-trained regression prediction model to generate a target model parameter in the plurality of sets of model parameters of different scales that meets a preset real-time perception performance index requirement; a second generation unit configured to, after a preset time interval, update the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, and the AP index according to a vehicle-side perception model and a road-side perception model corresponding to the target model parameter to obtain a new vehicle-cloud transmission time delay, a new road-cloud transmission time delay, a new environment feature vector, and a new AP index corresponding to a second frame of the multi-frame cooperative perception result, and input the new vehicle-cloud transmission time delay, the new road-cloud transmission time delay, the new environment feature vector, the new AP index, and the target model parameter into the regression prediction model to generate a new target model parameter in the plurality of sets of model parameters of different scales that meets the preset real-time perception performance index requirement; and an iteration unit configured to iteratively perform an update operation of the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, the AP index, and the target model parameter until a target model parameter corresponding to each frame of the multi-frame cooperative perception result is generated, so as to obtain a target optimized vehicle-side perception model and a target optimized road-side perception model corresponding to each frame of the multi-frame cooperative perception result based on the target model parameter corresponding to each frame of the multi-frame cooperative perception result.

[0020] The third aspect of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the vehicle-road cloud real-time cooperative perception method for an intelligent network environment as described in the above embodiments.

[0021] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and the program is executed by a processor to implement the vehicle-road cloud real-time cooperative perception method for an intelligent network environment as described above.

[0022] The fifth aspect embodiment of the present application provides a computer program product comprising a computer program executed to implement the intelligent network environment-oriented vehicle-road cloud real-time collaborative perception method described above.

[0023] Therefore, the embodiments of the present application have the following beneficial effects:

[0024] The embodiments of the present application can obtain multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene respectively, and input the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into a preset vehicle-end perception model and a road-end perception model respectively to generate vehicle-end perception results and road-end perception results; the vehicle-end perception results and the road-end perception results are fused at a target edge cloud end to obtain a collaborative perception result and a fusion time length, and vehicle-to-cloud transmission time delays and road-to-cloud transmission time delays between at least one target vehicle-end information node and at least one road-end information node and the target edge cloud end are obtained respectively, so as to calculate a total fusion output time length based on the vehicle-to-cloud transmission time delays, the road-to-cloud transmission time delays and the fusion time length, and calculate a real-time perception performance index corresponding to the collaborative perception result according to the total fusion output time length and a preset perception output lag parameter; the vehicle-end perception model and the road-end perception model are subjected to a preset dynamic optimization operation based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and the vehicle-road cloud real-time collaborative perception result is generated according to the target optimized vehicle-end perception model and the target optimized road-end perception model. The real-time accurate performance of the target-level collaborative perception algorithm can be improved by the dynamic adjustment strategy. Therefore, the problem that the prior art cannot model the communication time delay of the perception information transmission to the cloud in the intelligent networked automobile cloud control system from the algorithm perspective, and the mutual restriction of the calculation error and the time delay of the collaborative perception model, and the need to use different scale models for different environmental complexity are solved.

[0025] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter in the description of the application. BRIEF DESCRIPTION OF DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the accompanying drawings, in which:

[0027] Figure 1 A flowchart of an intelligent network environment-oriented vehicle-road cloud real-time collaborative perception method according to an embodiment of the present application is provided.

[0028] Figure 2 A vehicle-road cloud collaborative perception real-time algorithm framework diagram according to an embodiment of the present application is provided.

[0029] Figure 3 A schematic diagram of a target-level vehicle-road-cloud collaborative target detection algorithm framework provided for one embodiment of the present application;

[0030] Figure 4 A schematic diagram of the execution logic of a point cloud-based object-level vehicle-road-cloud collaborative target detection algorithm provided by one embodiment of the present application;

[0031] Figure 5 A schematic diagram of the calculation process of a vehicle-road-cloud collaborative perception algorithm provided for one embodiment of the present application;

[0032] Figure 6 A schematic diagram of collaborative sensing calculation results under extreme latency conditions provided by an embodiment of the present application;

[0033] Figure 7 A schematic diagram of the logical architecture of a dynamic optimization strategy provided for one embodiment of the present application;

[0034] Figure 8 An embodiment of the present application provides a f conservative With f globalbest Schematic diagram of algorithm operation comparison;

[0035] Figure 9 This is an example diagram of a vehicle-road-cloud real-time collaborative perception device for an intelligent connected environment according to an embodiment of the present application;

[0036] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0037] Among them, 10-real-time vehicle-road-cloud collaborative perception device for intelligent connected environment; 100-perception module, 200-indicator construction module, 300-dynamic optimization module; 1001-memory, 1002-processor, 1003-communication interface. DETAILED DESCRIPTION

[0038] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0039] A method and device for real-time collaborative perception of vehicle-road cloud in an intelligent network environment are described below with reference to the accompanying drawings. In view of the problems mentioned in the background art, the present application provides a method for real-time collaborative perception of vehicle-road cloud in an intelligent network environment. In this method, multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene are obtained respectively, and the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data are input into a preset vehicle-end perception model and a road-end perception model respectively to generate vehicle-end perception results and road-end perception results. The vehicle-end perception results and the road-end perception results are fused at a target edge cloud to obtain a collaborative perception result and a fusion time length. The vehicle-cloud transmission delay and the road-cloud transmission delay between the at least one target vehicle-end information node, the at least one road-end information node, and the target edge cloud are obtained respectively. Based on the vehicle-cloud transmission delay, the road-cloud transmission delay, and the fusion time length, the total fusion output time length is calculated, and the real-time perception performance index corresponding to the collaborative perception result is calculated according to the total fusion output time length and a preset perception output lag parameter. Based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy, a preset dynamic optimization operation is performed on the vehicle-end perception model and the road-end perception model to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and a vehicle-road cloud real-time collaborative perception result is generated according to the target optimized vehicle-end perception model and the target optimized road-end perception model. The real-time accurate performance of the target level collaborative perception algorithm can be improved by dynamically adjusting the strategy. Thus, the problem that the prior art cannot model the communication delay of the perception information transmission to the cloud in the intelligent networked vehicle cloud control system from the algorithm perspective, and the mutual restriction of calculation error and delay of the collaborative perception model, and the need to use different scale models for different environmental complexity are solved.

[0040] Specifically, Figure 1 A flowchart of a method for real-time collaborative perception of vehicle-road cloud in an intelligent network environment is provided in the embodiments of the present application.

[0041] As Figure 1 shown, the method for real-time collaborative perception of vehicle-road cloud in an intelligent network environment includes the following steps:

[0042] In step S101, multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene are obtained respectively, and the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data are input into a preset vehicle-end perception model and a road-end perception model respectively to generate vehicle-end perception results and road-end perception results.

[0043] It can be understood that in the embodiments of the present application, vehicle-road-cloud collaborative perception is a subtask of the overall perception task of the intelligent connected vehicle cloud control system in an intelligent connected environment. Its main function is to use information from different perspectives on both sides of the vehicle and road to perform collaborative perception of traffic scenes, thereby helping to achieve the task of generating digital twins and real-time mapping in the edge cloud.

[0044] It should be noted that the real-time algorithm framework of vehicle-road-cloud collaborative perception in the embodiment of this application mainly involves the vehicle side, road side and edge cloud side, such as Figure 2 As shown, the embodiment of the present application has a specific definition of modular information transmission and algorithm execution process based on the respective characteristics of the three ends of the vehicle, road, and cloud. Specifically, due to the characteristics of wireless transmission between the vehicle and the edge cloud, when raw-level information with high transmission load requirements cannot be transmitted between the vehicle and the edge cloud, the embodiment of the present application can transmit target-level information between the vehicle and the edge cloud, as well as feature-level information with lower load requirements after encoding and information screening. Since the information communication between the road and the edge cloud is transmitted through a wired connection with a large load-bearing capacity, the embodiment of the present application can be defined as a transmission that can carry raw-level data.

[0045] Therefore, the embodiment of the present application can respectively obtain multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to the vehicle-end information nodes and road-end information nodes in the traffic scene, and input them into the vehicle-end perception model and the road-end perception model respectively, thereby generating vehicle-end perception results and road-end perception results.

[0046] Optionally, in one embodiment of the present application, multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in the target traffic scene are respectively obtained, and the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data are respectively input into the preset vehicle-end perception model and the road-end perception model to generate vehicle-end perception results and road-end perception results, including: obtaining multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data respectively through at least one target vehicle-end information node and at least one road-end information node, and using the vehicle-end perception model to detect and track the multi-frame vehicle-end point cloud data to generate vehicle-end perception results; using the road-end perception model to perform road-end target detection and feature extraction operations on the multi-frame road-end point cloud data to generate road-end perception results.

[0047] In the actual implementation process, the operation process of the vehicle-road-cloud collaborative perception real-time algorithm in the embodiment of the present application on the vehicle side is as follows:

[0048] 1. K-frame vehicle endpoint cloud data acquisition;

[0049] 2. Fusion target detection and tracking: using target detection and tracking algorithms (i.e., vehicle-end perception model) to detect and track K-frame data, thereby combining the target-level data transmitted by the edge cloud in the previous frame to perform trajectory matching and updating, to generate vehicle-end perception results;

[0050] 3. Region of interest determination: determining the region of interest based on trajectory rules and the vehicle-end blind spot filling field of view information;

[0051] 4. Region of interest encoding: using feature extractors and feature selectors to encode the region of interest and upload it to the edge cloud;

[0052] 5. Obtain K+1 frame vehicle-end point cloud data.

[0053] The execution process of the vehicle-road cloud collaborative perception real-time algorithm in the edge cloud is as follows:

[0054] 1. Fusion target detection and feature extraction: extract the perception information features of the road end;

[0055] 2. Vehicle-road target-level information fusion detection and tracking: fuse vehicle-road information data to achieve preliminary ID consistent continuous tracking;

[0056] 3. Information decoding: decode the feature-level information uploaded by the vehicle end to obtain feature information;

[0057] 4. Region feature fusion: fuse vehicle-road information region feature-level data;

[0058] 5. Target detection and inference tracking: region of interest fusion detection and tracking, issue fusion tracking results and assist in updating the edge cloud digital twin model.

[0059] The roadside node can obtain multiple frames of road-end point cloud data, and perform road-end target detection and feature extraction operations on the multiple frames of road-end point cloud data through the road-end perception model to generate road-end perception results.

[0060] It should be noted that the embodiments of the present application can be based on the vehicle-road cloud collaborative perception algorithm framework, and design an algorithm framework for vehicle-road cloud collaborative target detection algorithm using target-level perception algorithm, which is as shown in Figure 2 The specific algorithm framework is as shown in Figure 3 The vehicle-road cloud collaborative perception algorithm framework is simplified, and mainly focuses on the target-level perception algorithm on both sides of the vehicle and road and the target-level transmission of information.

[0061] Therefore, by designing the vehicle-road cloud collaborative perception real-time algorithm framework, combining the vehicle-road cloud communication and perception characteristics, and considering the communication delay and real-time perception requirements, the embodiments of the present application design a target-level perception architecture for the intelligent connected vehicle cloud control system that adapts to the delay compensation mechanism, and effectively achieve real-time, efficient and stable traffic perception domain perception result acquisition.

[0062] In step S102, the vehicle-side perception results and the road-side perception results are fused through the target edge cloud to obtain the collaborative perception results and fusion duration, and the vehicle-cloud transmission delay and road-cloud transmission delay between at least one target vehicle-side information node and at least one road-side information node and the target edge cloud are respectively obtained, so as to calculate the total fusion output duration based on the vehicle-cloud transmission delay, the road-cloud transmission delay and the fusion duration, and calculate the real-time perception performance index corresponding to the collaborative perception result based on the total fusion output duration and the preset perception output lag parameter.

[0063] Furthermore, embodiments of the present application also require fusing the vehicle-side perception results and road-side perception results through the edge cloud to obtain collaborative perception results and obtain the fusion duration. Afterwards, embodiments of the present application can respectively obtain the vehicle-to-cloud transmission delay and road-to-cloud transmission delay between the vehicle-side information node, the road-side information node, and the target edge cloud, thereby calculating the total fusion output duration using the vehicle-to-cloud transmission delay, the road-to-cloud transmission delay, and the fusion duration, and calculating the real-time perception performance indicator PAP (Practical Average Precision) based on the total fusion output duration and the perception output lag parameter. Define the evaluation indicators for the real-time algorithm of vehicle-road-cloud collaborative perception.

[0064] Therefore, the embodiment of the present application determines the real-time perception performance indicators of vehicle-road-cloud collaborative perception, thereby modeling the impact of vehicle-road-cloud communication delay and computing delay on the real-time performance of the collaborative perception algorithm.

[0065] Optionally, in one embodiment of the present application, the vehicle-side perception results and the road-side perception results are fused through the target edge cloud to obtain collaborative perception results and fusion duration, including: performing coordinate conversion processing on the vehicle-side perception results and the road-side perception results through the target edge cloud, and performing target-level information fusion detection operations on the vehicle-side perception results and the road-side perception results after the coordinate conversion processing to generate collaborative perception results and fusion duration.

[0066] It should be noted that in the edge cloud fusion process, the embodiment of the present application can use point cloud data as input at both ends, such as Figure 4 As shown, the point cloud data at both ends of the vehicle and the road are input into their respective target detection algorithms (i.e., the vehicle-side perception model and the road-side perception model) to obtain corresponding vehicle-road target-level results (i.e., the vehicle-side perception results and the road-side perception results); thereafter, the embodiment of the present application can upload the vehicle-road target-level results to the cloud, and perform coordinate conversion and alignment on the cloud, and after target-level fusion, output the collaborative perception results, and at the same time calculate the fusion time of the fusion process on the edge cloud.

[0067] Therefore, the embodiment of the present application provides reliable data guidance and basis for the realization of the vehicle-road cloud real-time cooperative perception by acquiring the cooperative perception result and the fusion duration of the edge cloud.

[0068] Optionally, in an embodiment of the present application, the vehicle-cloud transmission delay and the road-cloud transmission delay between the at least one target vehicle-end information node and the at least one road-end information node and the target edge cloud are respectively acquired, the total fusion output duration is calculated based on the vehicle-cloud transmission delay, the road-cloud transmission delay and the fusion duration, and the real-time perception performance index corresponding to the cooperative perception result is calculated according to the total fusion output duration and the preset perception output lag parameter, including: determining the road-side model perception duration corresponding to the vehicle-end perception model and the vehicle-side model perception duration corresponding to the road-end perception model; calculating the total fusion output duration according to the road-side model perception duration, the vehicle-side model perception duration, the vehicle-cloud transmission delay, the road-cloud transmission delay and the fusion duration; comparing the total fusion output duration with the perception output lag parameter; when the total fusion output duration is less than the perception output lag parameter, calculating the first perception result performance corresponding to each frame of vehicle-end point cloud data in the multiple frames of vehicle-end point cloud data; when the total fusion output duration is greater than or equal to the perception output lag parameter, calculating the second perception result performance corresponding to each frame of vehicle-end point cloud data in the multiple frames of vehicle-end point cloud data; acquiring the total number of frames of the multiple frames of vehicle-end point cloud data, and calculating the total real-time perception performance corresponding to the multiple frames of vehicle-end point cloud data according to the first perception result performance and the second perception result performance, and calculating the real-time perception performance index based on the total real-time perception performance and the total number of frames.

[0069] For the establishment of the real-time perception performance index in the vehicle-road cloud cooperative perception, the present embodiment first needs to determine the source of the perception error in the intelligent network connection environment, which can usually use the number of perception data processed per second to describe the high real-time of input, and use the frame rate FPS (Frame Per Second) to represent, whose unit is Hz.

[0070] It should be noted that the target-level cooperative perception algorithm calculation process in the embodiment of the present application is as shown in Figure 5 (t 0+α -t0) represents the calculation duration of the target-level data output by the target-level perception algorithm loaded on the respective computing platforms for the data input at t0 at the vehicle and road ends; (t 0+α+β -t 0+α ) is the matching delay, that is, the target-level data transmission duration of the target-level perception result output by the target-level perception algorithm at the vehicle and road ends in the target-level data transmission process in the target-level perception target matching domain association process on the edge cloud, and the above two calculation durations together constitute the delay of the target-level vehicle-road cloud cooperative perception algorithm as (t 0+α+β -t0).

[0071] In the embodiment of the present application, as shown in Figure 5The perception results on the right are shown in the legend. The red truck input in the t0 vehicle-road data pair is used as the collaborative perception example target. The blue target detection box is the perception output of the collaborative perception algorithm for the input data at time t0. The classification results are The perception output result of the input data at time t0 is shown in the following formula:

[0072]

[0073] in, Respectively represent the geometric center point of the perception results of this truck in the x, y, and z directions; Represents the perceived size of the truck in the x, y, and z directions respectively.

[0074] Similarly, in the embodiment of the present application, the red target detection frame can be set as the target detection frame when the collaborative perception algorithm is at t 0+α+β Time output The corresponding current truck type is And real target information As shown in the following formula:

[0075]

[0076] It is understandable that Figure 5 The gap between the target's true position, represented by the purple arrow, and the perception result output by the perception algorithm, i.e., the perception error caused by the time delay, can be expressed as Delays may also cause and Any discrepancy between the two is a false positive.

[0077] At the same time, if the delay is too large, it may cause Figure 6 The results shown are as follows Figure 6 The output of the collaborative perception algorithm shown in the blue box does not overlap with the red box representing the actual position of the current target, which may lead to the situation where the perception results cannot be used and cannot be matched.

[0078] The following embodiments of the present application will provide a formal description and definition of the real-time collaborative perception problem in the research scenario described above.

[0079] It should be noted that the observation data stream in the embodiment of the present application can be modeled as a set of sensor observations, that is, a set of a series of world states and timestamps. For the 3D object detection problem, it can be expressed as:

[0080]

[0081] Among them, X jis the original data of point cloud; Y j is the ground truth of X j is the ground truth of X j is the ground truth of X j is the time stamp of X.

[0082] For an original data X with N targets, the ground truth Y can be represented as:

[0083]

[0084] where, is the geometric center of the target in the point cloud original data, and are the length, width and height of the target in the point cloud original data, respectively. i is the category of the target.

[0085] Therefore, for vehicle-road cloud collaborative perception, the point cloud data on both sides of the vehicle and the road can be represented as and which represent the data flow of the vehicle end and the road end, respectively; meanwhile, since there is a transmission delay process from the vehicle to the cloud and a transmission delay process from the road to the cloud in this problem, the two transmission delay processes can be defined as and wherein the former is the transmission delay from the vehicle to the cloud, and the latter is the transmission delay from the road to the cloud.

[0086] In a typical vehicle-road cloud collaborative perception scenario, for each frame of sensor perception data of the vehicle and the road, the comprehensive time from obtaining environmental perception data to outputting the fused perception result on the edge cloud can be represented as T c , as shown in the following formula:

[0087] T c = max {(T fv + T v ), (T fr + T r )} + T ff (5)

[0088] wherein T fv represents the time consumption of the vehicle end target detection algorithm; T fr represents the time consumption of the road end target detection algorithm; T v represents the data transmission time from the vehicle to the edge cloud; T r represents the transmission time from the road end node to the cloud; and T ffRepresents the time consumption of the fusion algorithm. The time consumption component of this process means that in the research scenario, the process of the vehicle performing perception calculation and transmitting the perception results to the cloud is in parallel with the process of the road performing perception calculation and transmitting the perception results to the cloud. The edge cloud can collect data from vehicles and roadsides and perform fusion perception.

[0089] The specific process of real-time collaborative perception can be defined as: at the real value demand output time t′ N The output of the previous vehicle-road-cloud collaborative perception algorithm based on transmission delay, which uses past observation data from the vehicle-side and roadside nodes as input, is as follows:

[0090]

[0091] The algorithm perception result of the Nth frame is defined as:

[0092]

[0093] The timestamp of the perception output is t output , which means that the algorithm can use time t N And all previous data frames are calculated, where t N =t′ N -ε, ε is the perception output hysteresis parameter set artificially according to the algorithm delay, transmission delay and deployment requirements.

[0094] It should be noted that t output ≤t′ N , and in most cases t output ≠t′ N , because the result of the collaborative sensing algorithm Can be at t′ N Any time output before is acceptable; meanwhile, the result Timestamp t output Less than t′ N The closest timestamp.

[0095] Based on the above-described vehicle-road-cloud collaborative real-time perception problem, the performance evaluation of the vehicle-road-cloud collaborative real-time perception problem is defined as the time t N And all previous data frames are calculated, and the required perception truth value output time t′ N The perception results output before Therefore, for the Nth output, the embodiment of the present application can compare and Y N Evaluate the performance of collaborative sensing algorithms.

[0096] The embodiment of the present application is based on the above process for the real-time collaborative perception problem, so that a new evaluation index PAP can be defined. The pseudo code of its specific evaluation is shown in the following pseudo code:

[0097]

[0098]

[0099] It should be understood by those skilled in the art that the traditional evaluation metric AP (Average Precision) cannot be used directly to evaluate the performance of this real-time problem, because the true value and predicted value of the data frame involved in the calculation in the AP value evaluation method are the same frame in all cases. Specifically, for t i The original data X at the moment i Prediction data And the true value data Y i For AP calculation, the data pairs involved are

[0100] For the real-time collaborative perception problem, due to the limitation of perception output time ε, i The original data at the moment; for the original data X i Prediction data And the true value data Y i In terms of, if Y i Output time point t i+α Greater than t i +ε, then this frame data cannot be i The perception output node corresponding to time t i +ε is used. However, due to the demand of downstream applications, such as digital twins, for continuous perception output, the current time point t i +εThe latest perception data available is the one using t i-1 Original data X i-1 Output prediction data Therefore, for the collaborative perception problem, the data that needs to be calculated in real time for the accuracy performance in this embodiment of the application should be

[0101] Therefore, the embodiment of the present application expands the offline performance indicator AP in the real-time perception problem into PAP through the above process, thereby characterizing the accurate performance of the perception results that can actually be output at the perception demand output node.

[0102] In step S103, based on the real-time perception performance index and the preset real-time dynamic algorithm optimization strategy, the vehicle-end perception model and the road-end perception model are subjected to preset dynamic optimization operation to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and a vehicle-road cloud real-time collaborative perception result is generated according to the target optimized vehicle-end perception model and the target optimized road-end perception model.

[0103] Further, the embodiment of the present application can determine a model dynamic optimization method of a real-time collaborative target detection algorithm suitable for a smart connected vehicle cloud control system based on the collaborative perception architecture and the defined perception real-time algorithm evaluation index, so as to realize real-time accurate perception of traffic vehicles in a traffic domain.

[0104] Optionally, in an embodiment of the present application, based on the real-time perception performance index and the preset real-time dynamic algorithm optimization strategy, the vehicle-end perception model and the road-end perception model are subjected to preset dynamic optimization operation to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, including: performing feature engineering environment feature modeling on a target traffic scene corresponding to each frame of collaborative perception result in multiple frames of collaborative perception results to obtain an environment feature vector corresponding to the target traffic scene; constructing multiple groups of model parameter sets of different scales for the vehicle-end perception model and the road-end perception model respectively, and determining an AP index corresponding to each group of model parameter sets of different scales in the multiple groups of model parameter sets of different scales; inputting a vehicle-cloud transmission time delay, a road-cloud transmission time delay, an environment feature vector and an AP index corresponding to a first frame of collaborative perception result in the multiple frames of collaborative perception results into a pre-trained regression prediction model to generate target model parameters in the multiple groups of model parameter sets of different scales that meet the preset real-time perception performance index requirements; after a preset time interval, updating the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector and the AP index according to the vehicle-end perception model and the road-end perception model corresponding to the target model parameters to obtain new vehicle-cloud transmission time delay, new road-cloud transmission time delay, new environment feature vector and new AP index corresponding to a second frame of collaborative perception result, and inputting the new vehicle-cloud transmission time delay, the new road-cloud transmission time delay, the new environment feature vector, the new AP index and the target model parameters into the regression prediction model to generate new target model parameters in the multiple groups of model parameter sets of different scales that meet the preset real-time perception performance index requirements; iteratively performing the updating operation of the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, the AP index and the target model parameters until the target model parameters corresponding to each frame of collaborative perception result are generated, so as to obtain the target optimized vehicle-end perception model and the target optimized road-end perception model corresponding to each frame of collaborative perception result based on the target model parameters corresponding to each frame of collaborative perception result.

[0105] Specifically, in the embodiments of the present application, the main computing delay of the target level cooperative algorithm framework exists in the point cloud target detection model at both ends of the vehicle-road, and the embodiments of the present application can set an optimization model here to dynamically optimize the point cloud target detection algorithm, such as Figure 7 As shown in the red line on the left, this is the specific running process of the dynamic vehicle-road model for the optimization strategy. In addition, Figure 7 The red line on the right is the research process of the dynamic optimization strategy of the dynamic vehicle-road model.

[0106] Based on the above modeling of the time process of vehicle-road cloud cooperative perception, the following is obtained:

[0107] T c = max {(T fv + T v ), (T fr + T r )} + T ff (8)

[0108] It can be understood that the best algorithm flow is to select the best algorithm parameter M with the best performance as the best solution through the offline training process under the assumption that T fv , T v , T fr , T r and T ff are fixed, so that T c is also fixed. The algorithm in this case is called f globalbest , which cannot be dynamically adjusted according to different perception and communication environments, and may not be able to complete the perception result output of t′ N before the data input of t N . To solve this problem, the algorithm f conservative can be inclined to use the model parameter pair with less time consumption when deployed, and the running comparison process of the two is as shown in Figure 8

[0109] In actual execution process, the embodiments of the present application can update the parameters M of the vehicle-side and roadside detection algorithms online and dynamically at a time interval of Δt during deployment, so that the entire vehicle-road cooperative perception algorithm can adapt to and dynamically adjust to the current perception situation, thereby achieving optimal performance. The flow is f dynamic , as shown in the following pseudo code:

[0110]

[0111]

[0112] ​To formulate the above dynamic strategy, the embodiments of the present application need to consider establishing a sufficient representation of the current vehicle-road cloud collaborative perception situation to determine how to design the conversion strategy of the embodiments of the present application. Specifically, the factors affecting the current vehicle-road cloud collaborative perception situation mainly include the complexity of the perception environment e t , the model complexity caused by the model parameter set (M v , M r ), and the target level information transmission delay situation (T v , T r ) between the vehicle and the road cloud.

[0113] In the specific implementation process, the embodiments of the present application can predict the PAP of different model parameter sets (M v , M r ) under the given environment feature vector e t and transmission delay (T v , T r ), which actually converts the problem of representing the perception influencing factors into a regression problem. The embodiments of the present application train seven Pointpillars models with the same structure but different scales for vehicles and roads, and set the optimal model parameters (M vbest , M rbest ), and then select the optimal model parameter set used in the future Δt time.

[0114] As an implementable way, the embodiments of the present application train a random forest model to predict the parameters (M vbest , M rbest ), and use the mean square error as the loss function, and the regression model training process is shown in the following pseudo code:

[0115]

[0116]

[0117] In addition, the embodiments of the present application also need to supplement the following contents to the above pseudo code:

[0118] 1. Environment feature vector setting.

[0119] In order to fully improve the performance of the regression model and minimize the inference time, the embodiments of the present application perform feature engineering on the traffic scene based on the laser radar traffic data to model the environment features, and these features include:

[0120] 1) num_bbox (number of bounding boxes);

[0121] 2) bounding box occlusion (bounding box occlusion situation);

[0122] 3) bounding box truncation situation;

[0123] 4) littlepoints objects;

[0124] 5) time of day.

[0125] 2, parameter set of different scale Pointpillars (PP) model.

[0126] The size of the backbone network and the neck network of the basic model, i.e. the size of PP3, is shown in Table 1:

[0127] Table 1

[0128]

[0129] The embodiments of the present application can perform size adjustment of the model based on the depth scaling coefficient a≈1.23, the width scaling coefficient β≈1.61 and the scaling φ. In the size adjustment, the size of the convolution layer of each stage of the model is adjusted, and the AP indicators of 3D and BEV of different models are calculated to reflect their performance, as shown in Table 2:

[0130] Table 2

[0131]

[0132]

[0133] In the embodiments of the present application, the model calculation delay is used as the prior input of the algorithm execution model parameter search, and in the inference process, it can be used as the prior experience to execute the model parameter search. The embodiments of the present application collect the algorithm calculation delay of the algorithm model with different parameters in a large number of scenes to obtain the statistical data of the algorithm calculation delay, and use the statistical time data t average , t best , t worst as the input of the cooperative perception dynamic optimization algorithm f dynamic , as shown in Table 3:

[0134] Table 3

[0135]

[0136] It should be noted that the point cloud target detection data part of the existing Kitti data set can be used as a large-scale pre-training data set of the vehicle-road end multi-scale Pointpillar target detection model in the embodiments of the present application, and the vehicle-road point cloud target detection data part of the DAIR-V2X data set can be used as the main experimental verification and training data set.

[0137] For the data set used by the regression model involved in the dynamic optimization strategy, the embodiments of the present application can generate a new data set by using the frame-by-frame performance indicators generated after inference on the filtered DAIR-V2X data set based on the trained multi-scale point cloud target detection model, and the environment indicators generated by the above method, and combing the time sequence to generate training set and validation set respectively.

[0138] The PAP indicators and dynamic optimization strategies involved in the intelligent network environment-oriented vehicle-road cloud real-time collaborative perception method of the present application are quantitatively verified as follows.

[0139] 1. Quantitative verification of PAP indicators

[0140] The embodiments of the present application can select three groups of vehicle-road model pairs to prove that the performance of the model in the real-time perception scene may decrease compared to offline testing, and the three groups of vehicle-road model pairs can be (PP4, PP2), (PP6, PP4) and (PP7, PP6); the initial transmission delay T v0 is 50ms, T r0 is 20ms, and the fusion delay ε is 100ms. The AP and PAP values of the selected vehicle-road model pairs are shown in Table 4:

[0141] Table 4

[0142]

[0143] As can be seen from Table 4, the failure rate of the module pair with higher computational complexity is also larger in the real-time environment, resulting in a significant decrease in real-time accuracy performance. Finally, the module pair (PP7, PP6) performs worst among the three module pairs; at the same time, it is shown that the AP as a vector reflecting the offline accuracy of the model is significantly different from the PAP defined in the embodiments of the present application for real-time problems, and the PAP is a comprehensive index that takes into account real-time and accuracy.

[0144] 2. Quantitative verification of dynamic optimization strategy

[0145] In order to demonstrate the universality of the dynamic optimization strategy, the embodiments of the present application set the initial transmission delay T v0 to 50ms, 75ms and 100ms, and T r0The crossover experiment is conducted with 20ms, 35ms and 50ms and fusion delay ε is 100ms and 200ms. Table 5 shows the dynamic strategy algorithm f proposed in the embodiment of the present application. dynamic , global optimal single algorithm f globalbest and the conservative method with time redundancy f conservative performance.

[0146] Table 5

[0147]

[0148]

[0149] In the conservative deployment algorithm, the embodiment of the present application may choose to leave 50% of the time redundancy or an artificially set conservative threshold t p = 20ms, so that it can be used after completing the vehicle and road perception processing, which means that the conservative strategy ensures that the output of each perception result can be used, that is, the failure rate is 0%; in terms of vehicle-cloud fusion, the embodiment of the present application can use the classic weighted Hungarian matching algorithm, and its calculation time per frame is on the order of 10 -2 ms, so it can be ignored.

[0150] Taking the first set of settings as an example, for f globalbest In most cases, PP6, which ranks second in accuracy when deployed on the roadside, and PP4, which ranks fourth in accuracy when deployed on the vehicle side, can achieve the highest perception accuracy while meeting the perception output requirements. However, due to fluctuations in the perception environment and transmission delays, the results cannot be fused in time in some frames, resulting in lower perception accuracy in some frames. globalbest This performance is lost in these frames, thus providing room for perceptual optimization in these frames.

[0151] The dynamic algorithm in the embodiment of the present application is most of the time the same as the global optimal single algorithm f globalbest However, in frames where the global optimal algorithm cannot output perception results, it can output perception results in real time by dynamically adjusting model parameters, thereby improving the overall perception performance under the constraints of perception output requirements; at the same time, compared with the conservatively deployed algorithm, the excess redundant time of model deployment is fully utilized, achieving higher accuracy, which improves dynamic perception performance while ensuring output security.

[0152] After sufficient data verification, the algorithm f designed in this application embodiment for dynamic adjustment of computing scenarios dynamic , and f globalbest There is a significant dynamic PAP improvement of about 5.8% compared to f conservativeCompared with the increase of 27.5%, thus proving the necessity of dynamically adjusting the parameter M and f dynamic The effectiveness of the design.

[0153] The vehicle-road cloud real-time collaborative perception method for an intelligent network environment according to the embodiment of the application, by respectively acquiring multi-frame vehicle end point cloud data and multi-frame road end point cloud data corresponding to at least one target vehicle end information node and at least one road end information node in a target traffic scene, and inputting the multi-frame vehicle end point cloud data and the multi-frame road end point cloud data into a preset vehicle end perception model and a road end perception model respectively, to generate vehicle end perception results and road end perception results; by fusing the vehicle end perception results and the road end perception results at the target edge cloud end to obtain a collaborative perception result and a fusion time length, and respectively acquiring a vehicle cloud transmission time delay and a road cloud transmission time delay between at least one target vehicle end information node and at least one road end information node and the target edge cloud end, to calculate a total fusion output time length based on the vehicle cloud transmission time delay, the road cloud transmission time delay and the fusion time length, and calculate a real-time perception performance index corresponding to the collaborative perception result according to the total fusion output time length and a preset perception output lag parameter; based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy, performing a preset dynamic optimization operation on the vehicle end perception model and the road end perception model to obtain a target optimized vehicle end perception model and a target optimized road end perception model, and generating a vehicle-road cloud real-time collaborative perception result according to the target optimized vehicle end perception model and the target optimized road end perception model. The application can improve the real-time accuracy performance of the target level collaborative perception algorithm through a dynamic adjustment strategy.

[0154] Secondly, referring to the accompanying drawings, a vehicle-road cloud real-time collaborative perception device for an intelligent network environment according to the embodiment of the application is described.

[0155] Figure 9 is a block schematic diagram of the vehicle-road cloud real-time collaborative perception device for an intelligent network environment according to the embodiment of the application.

[0156] As Figure 9 shown, the vehicle-road cloud real-time collaborative perception device for an intelligent network environment 10 includes a perception module 100, an index construction module 200 and a dynamic optimization module 300.

[0157] The perception module 100 is configured to respectively acquire multi-frame vehicle end point cloud data and multi-frame road end point cloud data corresponding to at least one target vehicle end information node and at least one road end information node in a target traffic scene, and input the multi-frame vehicle end point cloud data and the multi-frame road end point cloud data into a preset vehicle end perception model and a road end perception model respectively, to generate vehicle end perception results and road end perception results.

[0158] The index construction module 200 is configured to fuse the vehicle-end perception result and the road-end perception result through the target edge cloud to obtain a cooperative perception result and a fusion time length, and obtain a vehicle-to-cloud transmission time delay and a road-to-cloud transmission time delay between at least one target vehicle-end information node and at least one road-end information node and the target edge cloud, calculate a total fusion output time length based on the vehicle-to-cloud transmission time delay, the road-to-cloud transmission time delay and the fusion time length, and calculate a real-time perception performance index corresponding to the cooperative perception result according to the total fusion output time length and a preset perception output lag parameter.

[0159] The dynamic optimization module 300 is configured to perform a preset dynamic optimization operation on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy, to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and generate a vehicle-road cloud real-time cooperative perception result according to the target optimized vehicle-end perception model and the target optimized road-end perception model.

[0160] Optionally, in an embodiment of the present application, the perception module 100 comprises an acquisition unit and a feature extraction unit.

[0161] The acquisition unit is configured to acquire multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data from the at least one target vehicle-end information node and the at least one road-end information node respectively, and detect and track the multi-frame vehicle-end point cloud data by using the vehicle-end perception model to generate a vehicle-end perception result.

[0162] The feature extraction unit is configured to perform road-end target detection and feature extraction operation on the multi-frame road-end point cloud data by using the road-end perception model to generate a road-end perception result.

[0163] Optionally, in an embodiment of the present application, the index construction module 200 comprises a fusion unit configured to perform coordinate conversion processing on the vehicle-end perception result and the road-end perception result through the target edge cloud, and perform target-level information fusion detection operation on the vehicle-end perception result and the road-end perception result after the coordinate conversion processing to generate a cooperative perception result and a fusion time length.

[0164] Optionally, in an embodiment of the present application, the index construction module 200 further comprises a first determination unit, a first calculation unit, a comparison unit, a second calculation unit, a third calculation unit and a fourth calculation unit.

[0165] The first determination unit is configured to determine a road-side model perception time length corresponding to the vehicle-end perception model and a vehicle-side model perception time length corresponding to the road-end perception model.

[0166] The first calculation unit is configured to calculate a total fusion output time length according to the road-side model perception time length, the vehicle-side model perception time length, the vehicle-to-cloud transmission time delay, the road-to-cloud transmission time delay and the fusion time length.

[0167] a comparison unit configured to compare the fusion output total duration and the perception output lag parameter.

[0168] a second calculation unit configured to calculate a first perception result performance corresponding to each frame of vehicle-end point cloud data in the plurality of frames of vehicle-end point cloud data when the fusion output total duration is less than the perception output lag parameter.

[0169] a third calculation unit configured to calculate a second perception result performance corresponding to each frame of vehicle-end point cloud data in the plurality of frames of vehicle-end point cloud data when the fusion output total duration is greater than or equal to the perception output lag parameter.

[0170] a fourth calculation unit configured to obtain a total number of frames of the plurality of frames of vehicle-end point cloud data, and calculate a total real-time perception performance of the plurality of frames of vehicle-end point cloud data according to the first perception result performance and the second perception result performance, and calculate a real-time perception performance index based on the total real-time perception performance and the total number of frames.

[0171] Optionally, in an embodiment of the present application, the dynamic optimization module 300 comprises a modeling unit, a second determination unit, a first generation unit, a second generation unit, and an iteration unit.

[0172] The modeling unit is configured to model a target traffic scene corresponding to each frame of cooperative perception result in the plurality of frames of cooperative perception result by feature engineering environment features, to obtain an environment feature vector corresponding to the target traffic scene.

[0173] The second determination unit is configured to construct a plurality of sets of model parameters of different scales for the vehicle-end perception model and the road-end perception model respectively, and determine an AP index corresponding to each set of model parameters of different scales in the plurality of sets of model parameters of different scales.

[0174] The first generation unit is configured to input a vehicle-cloud transmission delay, a road-cloud transmission delay, the environment feature vector, and the AP index corresponding to a first frame of cooperative perception result in the plurality of frames of cooperative perception result into a pre-trained regression prediction model, to generate a target model parameter in the plurality of sets of model parameters of different scales that meets a preset real-time perception performance index requirement.

[0175] The second generation unit is configured to, after a preset time interval, update the vehicle-cloud transmission delay, the road-cloud transmission delay, the environment feature vector, and the AP index according to the target model parameter corresponding to the vehicle-end perception model and the road-end perception model, to obtain a new vehicle-cloud transmission delay, a new road-cloud transmission delay, a new environment feature vector, and a new AP index corresponding to a second frame of cooperative perception result, and input the new vehicle-cloud transmission delay, the new road-cloud transmission delay, the new environment feature vector, the new AP index, and the target model parameter into the regression prediction model, to generate a new target model parameter in the plurality of sets of model parameters of different scales that meets the preset real-time perception performance index requirement.

[0176] an iteration unit configured to iteratively perform an updating operation of the vehicle-cloud transmission delay, the road-cloud transmission delay, the environmental feature vector, the AP index, and the target model parameter, until a target model parameter corresponding to each frame of the cooperative perception result is generated, so as to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model corresponding to each frame of the cooperative perception result based on the target model parameter corresponding to each frame of the cooperative perception result.

[0177] It should be noted that the foregoing explanation and description of the vehicle-road cloud real-time cooperative perception method for the intelligent network environment embodiment also applies to the vehicle-road cloud real-time cooperative perception device for the intelligent network environment embodiment, which will not be described here again.

[0178] The vehicle-road cloud real-time cooperative perception device for the intelligent network environment according to the embodiment of the present application comprises a perception module configured to acquire multiple frames of vehicle-end point cloud data and multiple frames of road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene respectively, and input the multiple frames of vehicle-end point cloud data and the multiple frames of road-end point cloud data into a preset vehicle-end perception model and a road-end perception model respectively to generate a vehicle-end perception result and a road-end perception result; an index construction module configured to fuse the vehicle-end perception result and the road-end perception result through a target edge cloud to obtain a cooperative perception result and a fusion time length, acquire a vehicle-cloud transmission delay and a road-cloud transmission delay between the at least one target vehicle-end information node and the at least one road-end information node and the target edge cloud respectively, calculate a total fusion output time length based on the vehicle-cloud transmission delay, the road-cloud transmission delay, and the fusion time length, and calculate a real-time perception performance index corresponding to the cooperative perception result according to the total fusion output time length and a preset perception output lag parameter; and a dynamic optimization module configured to perform a preset dynamic optimization operation on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and generate a vehicle-road cloud real-time cooperative perception result according to the target optimized vehicle-end perception model and the target optimized road-end perception model. The real-time accurate performance of the target-level cooperative perception algorithm can be improved by the dynamic adjustment strategy.

[0179] Figure 10 The structure schematic diagram of the electronic device provided by the embodiment of the present application is shown. The electronic device can comprise:

[0180] The memory 1001, the processor 1002, and the computer program stored in the memory 1001 and executable on the processor 1002.

[0181] The processor 1002 implements the vehicle-road cloud real-time cooperative perception method for the intelligent network environment provided in the above embodiments when executing the program.

[0182] Further, the electronic device further comprises:

[0183] The communication interface 1003 is configured to communicate between the memory 1001 and the processor 1002.

[0184] The memory 1001 is configured to store a computer program executable on the processor 1002.

[0185] The memory 1001 can include a high-speed RAM memory, and can also include a non-volatile memory, for example, at least one disk memory.

[0186] If the memory 1001, the processor 1002 and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001 and the processor 1002 can be connected to each other through a bus and complete communication between each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 10 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0187] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can complete communication between each other through an internal interface.

[0188] The processor 1002 can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.

[0189] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the above-mentioned vehicle-road cloud real-time collaborative perception method for an intelligent network environment.

[0190] The embodiment of the present application further provides a computer program product comprising a computer program, which, when executed, is configured to implement the intelligent network connection environment-oriented vehicle-road cloud real-time collaborative perception method.

[0191] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or N embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples, without contradiction.

[0192] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0193] Any process or method descriptions in flow charts or otherwise described herein represent embodiments that can be understood as a sequence of steps of executable instructions for implementing custom logic or processes, and the scope of preferred embodiments of the present application includes additional implementations that can not be precisely shown or discussed, including implementations that can be performed substantially concurrently, in reverse order, or in any other suitable order, according to the specific requirements of the function involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0194] The logic and / or steps represented in the flowcharts and / or described herein, for example, can be considered as a sequence of instructions to implement logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, processor- based system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a computer- readable storage medium or a computer-readable signal medium. The computer- readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires (electrical connections), a portable computer diskette (a magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, via, for example, optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.

[0195] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware and in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0196] Those skilled in the art can understand that all or part of the steps carried out by the above-mentioned embodiments can be completed by programs instructing related hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, one or a combination of the steps of the method embodiments is included.

[0197] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can be physically present separately, or two or more units can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0198] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A method for real-time cooperative perception of vehicle-road cloud in an intelligent networking environment, characterized in that, The method comprises the following steps: Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; Fusing the vehicle-end perception results and the road-end perception results through a target edge cloud to obtain a cooperative perception result and a fusion time length, respectively acquiring vehicle-cloud transmission time delays and road-cloud transmission time delays between the at least one target vehicle-end information node and the at least one road-end information node and the target edge cloud, and based on the vehicle-cloud transmission time delays, the road-cloud transmission time delays and the fusion time length, calculating a total fusion output time length, and according to the total fusion output time length and a preset perception output lag parameter, calculating a real-time perception performance index corresponding to the cooperative perception result; Based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy, performing a preset dynamic optimization operation on the vehicle-end perception models and the road-end perception models to obtain target optimized vehicle-end perception models and target optimized road-end perception models, and generating a vehicle-road cloud real-time cooperative perception result according to the target optimized vehicle-end perception models and the target optimized road-end perception models.

2. The method of claim 1, wherein, The method comprises the following steps: Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; 3. The method of claim 2, wherein, Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; The method comprises the following steps:

4. The method of claim 3, wherein, Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; Respectively acquiring multi-frame vehicle-end point cloud data and multi-frame road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene, and inputting the multi-frame vehicle-end point cloud data and the multi-frame road-end point cloud data into preset vehicle-end perception models and road-end perception models respectively to generate vehicle-end perception results and road-end perception results; determine a road-side model perception time length corresponding to the vehicle-end perception model and a vehicle-side model perception time length corresponding to the road-end perception model; calculate the fusion output total time length according to the road-side model perception time length, the vehicle-side model perception time length, the vehicle-cloud transmission time delay, the road-cloud transmission time delay, and the fusion time length; compare the fusion output total time length with the perception output lag parameter; when the fusion output total time length is less than the perception output lag parameter, calculate a first perception result performance corresponding to each frame of vehicle-end point cloud data in the multiple frames of vehicle-end point cloud data; when the fusion output total time length is greater than or equal to the perception output lag parameter, calculate a second perception result performance corresponding to each frame of vehicle-end point cloud data in the multiple frames of vehicle-end point cloud data; obtain a total frame number of the multiple frames of vehicle-end point cloud data, and calculate a real-time perception performance sum corresponding to the multiple frames of vehicle-end point cloud data according to the first perception result performance and the second perception result performance, and calculate the real-time perception performance index based on the real-time perception performance sum and the total frame number.

5. The method of claim 1, wherein, perform a preset dynamic optimization operation on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, including: perform feature engineering environment feature modeling on a target traffic scene corresponding to each frame of collaborative perception result in the multiple frames of collaborative perception result to obtain an environment feature vector corresponding to the target traffic scene; construct multiple groups of model parameter sets with different scales for the vehicle-end perception model and the road-end perception model respectively, and determine an average accuracy index corresponding to each group of model parameter set with different scales in the multiple groups of model parameter sets with different scales; input the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, and the average accuracy index corresponding to a first frame of collaborative perception result in the multiple frames of collaborative perception result into a pre-trained regression prediction model to generate a target model parameter in the multiple groups of model parameter sets with different scales that meets a preset real-time perception performance index requirement; after a preset time interval, update the vehicle-cloud transmission time delay, the road-cloud transmission time delay, the environment feature vector, and the average accuracy index according to the vehicle-end perception model and the road-end perception model corresponding to the target model parameter to obtain a new vehicle-cloud transmission time delay, a new road-cloud transmission time delay, a new environment feature vector, and a new average accuracy index corresponding to a second frame of collaborative perception result, and input the new vehicle-cloud transmission time delay, the new road-cloud transmission time delay, the new environment feature vector, the new average accuracy index, and the target model parameter into the regression prediction model to generate a new target model parameter in the multiple groups of model parameter sets with different scales that meets the preset real-time perception performance index requirement; The updating operations of the vehicle-cloud transmission delay, the road-cloud transmission delay, the environment feature vector, the average accuracy index, and the target model parameter are iteratively performed until the target model parameter corresponding to each frame of the cooperative perception result is generated, so as to obtain the target optimized vehicle-end perception model and the target optimized road-end perception model corresponding to each frame of the cooperative perception result based on the target model parameter corresponding to each frame of the cooperative perception result.

6. A vehicle-road cloud real-time collaborative perception device for an intelligent network environment, characterized in that, Comprise: The perception module is used for acquiring multiple frames of vehicle-end point cloud data and multiple frames of road-end point cloud data corresponding to at least one target vehicle-end information node and at least one road-end information node in a target traffic scene respectively, and inputting the multiple frames of vehicle-end point cloud data and the multiple frames of road-end point cloud data into a preset vehicle-end perception model and a road-end perception model respectively to generate vehicle-end perception results and road-end perception results. The index construction module is used for fusing the vehicle-end perception results and the road-end perception results through a target edge cloud to obtain a cooperative perception result and a fusion time length, acquiring a vehicle-cloud transmission delay and a road-cloud transmission delay between the at least one target vehicle-end information node and the at least one road-end information node and the target edge cloud respectively, calculating a total fusion output time length based on the vehicle-cloud transmission delay, the road-cloud transmission delay, and the fusion time length, and calculating a real-time perception performance index corresponding to the cooperative perception result according to the total fusion output time length and a preset perception output lag parameter. The dynamic optimization module is used for performing a preset dynamic optimization operation on the vehicle-end perception model and the road-end perception model based on the real-time perception performance index and a preset real-time dynamic algorithm optimization strategy to obtain a target optimized vehicle-end perception model and a target optimized road-end perception model, and generating a vehicle-road cloud real-time cooperative perception result according to the target optimized vehicle-end perception model and the target optimized road-end perception model.

7. The apparatus of claim 6, wherein, The perception module comprises: The acquisition unit is used for acquiring multiple frames of vehicle-end point cloud data and multiple frames of road-end point cloud data through the at least one target vehicle-end information node and the at least one road-end information node respectively, and detecting and tracking the multiple frames of vehicle-end point cloud data by using the vehicle-end perception model to generate the vehicle-end perception results. The feature extraction unit is used for performing road-end target detection and feature extraction operations on the multiple frames of road-end point cloud data by using the road-end perception model to generate the road-end perception results.

8. An electronic device, comprising: Comprise: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the vehicle-road cloud real-time cooperative perception method for an intelligent network environment as claimed in any one of claims 1-5.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the vehicle-road cloud real-time cooperative perception method for an intelligent network environment as claimed in any one of claims 1-5.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed to implement the vehicle-road cloud real-time cooperative perception method for an intelligent network environment as claimed in any one of claims 1-5.

Citation Information

Patent Citations

  • End-edge-cloud vehicle road collaborative fusion sensing architecture and construction method thereof

    CN113743479A

  • Characteristic-level collaborative perception fusion method and system for vehicle-road collaboration

    CN115578709A