Vehicle end data management method based on reinforcement learning, electronic equipment and storage medium

Through multi-sensor data acquisition and multi-agent enhancement learning model, the vehicle data governance strategy is dynamically adjusted, and the efficiency and safety of vehicle-side data governance in complex environments is solved, independent learning and optimization are achieved, and the intelligence and safety of the vehicle are improved.

CN120448773APending Publication Date: 2025-08-08XIAN KAIHUA ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474609.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing vehicle-side data governance methods are difficult to adapt to dynamic changes in complex traffic environments, resulting in low data processing efficiency, unreasonable resource allocation, rigid privacy protection strategies, and inability to achieve high data transmission delay and redundant storage costs.

Method used

Multi-sensor data acquisition and state modeling are used, and a multi-agent enhanced learning model is constructed using asynchronous advantageous action evaluation algorithm. The global resource state of the vehicle is sensed through multi-dimensional state space, the data acquisition frequency, processing priority and storage path are dynamically adjusted, and data privacy level evaluation and encryption algorithm matching are carried out through artificial intelligence technology.

Benefits of technology

It realizes independent learning and optimization of vehicle data governance, reduces data processing delays, improves storage resource utilization, reduces the risk of sensitive data leakage, adapts to environmental changes and task requirements, and improves vehicle performance and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448773A_ABST
    Figure CN120448773A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle end data management method based on reinforcement learning, electronic equipment and a storage medium. The method comprises the following steps: acquiring real-time traffic environment data and vehicle state data of a vehicle to obtain vehicle end original data; performing feature extraction and dimension reduction on the original data of the vehicle end by using an artificial intelligence technology, and constructing a multi-dimensional state space; determining a current vehicle state, analyzing the current state through a multi-agent reinforcement learning model, and determining a target execution strategy corresponding to the current vehicle to generate a target action; and evaluating a target action effect based on a reward function, generating a feedback signal, inputting the feedback signal into the multi-agent reinforcement learning model, and optimizing model parameters. According to the invention, full life cycle management of data acquisition, processing, storage and transmission is covered, real-time processing is carried out on the acquired real-time vehicle end data through the multi-agent decision model, an execution decision is obtained, limitation of static rules is broken through, and autonomous learning and optimization of a data governance strategy are realized at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent connected vehicle data processing technology, and in particular to a vehicle-side data management method, electronic device and storage medium based on reinforcement learning. Background Art

[0002] With the rapid development of intelligent connected vehicle technology, the massive amount of data collected in real time by on-board sensors (such as cameras, radars, and lidars) has placed higher demands on data governance. Traditional data governance methods typically rely on preset rules or static models, which are difficult to adapt to the dynamic changes in data in complex traffic environments. They suffer from problems such as low data processing efficiency, irrational resource allocation, and rigid privacy protection strategies. For example, when fusing multi-sensor data in existing technologies, data screening and aggregation are performed only through fixed thresholds. The strategy cannot be dynamically adjusted based on real-time road conditions, computing load, and other factors, resulting in high data transmission delays and redundant storage costs.

[0003] The development of artificial intelligence (AI) technology has provided new insights into vehicle-based data governance. However, existing solutions are mostly limited to data classification and simple cleaning, failing to fully leverage dynamic decision-making algorithms like reinforcement learning to autonomously optimize data governance strategies. Integrating the sequential decision-making capabilities of reinforcement learning to build an intelligent governance framework covering the entire data collection, processing, storage, and transmission process has become a key issue in improving the efficiency and security of vehicle-based data utilization. Summary of the Invention

[0004] The present invention provides a vehicle-side data governance method, electronic device and storage medium based on reinforcement learning. Through real-time collection and state modeling of multi-sensor data, and the use of an asynchronous dominant action evaluation algorithm (A3C) to build a multi-agent decision-making model, autonomous learning and optimization of data governance strategies are achieved.

[0005] The present invention provides a vehicle-side data management method based on reinforcement learning, comprising:

[0006] The vehicle's real-time traffic environment data, vehicle status data, and computing resource data are collected through multiple sensors on the vehicle to obtain the vehicle's original data;

[0007] Use artificial intelligence technology to extract features and reduce the dimension of vehicle-side raw data, and construct a multi-dimensional state space that includes environmental parameters, system resources, and data attributes;

[0008] Based on the multi-dimensional state space, the current vehicle state is determined, and the current state is analyzed through the multi-agent reinforcement learning model to determine the target execution strategy corresponding to the current vehicle and generate the target action;

[0009] The effect of the target action is evaluated based on the reward function. Based on the evaluation results, a feedback signal is generated and input into the multi-agent reinforcement learning model to optimize the model parameters.

[0010] Preferably, in a vehicle-side data governance method based on reinforcement learning, the method further includes:

[0011] An asynchronous advantage action evaluation algorithm is used to build a multi-agent reinforcement learning model, including:

[0012] Step A: Define each sensor node, data processing module and storage unit as an intelligent agent;

[0013] Step B: Based on the Markov decision process, establish the data governance task model corresponding to each agent and initialize it;

[0014] Step C: Control multiple agents to interact with the environment on different threads or processors, and collect multi-dimensional state space data, execution actions, and reward data of each agent during the environment interaction process;

[0015] Step D: Based on the update rules of the asynchronous advantage action evaluation algorithm, the data governance task model corresponding to each agent is updated using the multi-dimensional state space data, the executed actions, and the reward data for the executed actions;

[0016] Step E: Repeat steps CD until the accumulated reward corresponding to the strategy corresponding to the agent reaches the maximum value, and end the training of the data governance task model corresponding to the agent.

[0017] Preferably, in a vehicle-side data governance method based on reinforcement learning, artificial intelligence technology is used to extract features and reduce the dimension of the vehicle-side raw data to construct a multidimensional state space containing environmental parameters, system resources, and data attributes, including:

[0018] Identify the vehicle-side raw data based on artificial intelligence technology and determine the data type corresponding to the vehicle-side raw data;

[0019] Based on the data type, artificial intelligence technology is used to select a neural network to obtain the optimal dimensionality reduction network corresponding to the original data on the vehicle side;

[0020] Based on the optimal neural extraction network, data feature extraction and dimensionality reduction are performed on the corresponding original vehicle-side data to obtain the low-dimensional features of the original vehicle-side data;

[0021] The low-dimensional features are directly spliced to obtain a state vector, and the state vector is spatially mapped to generate a multi-dimensional state space.

[0022] Preferably, in a vehicle-side data governance method based on reinforcement learning, before using artificial intelligence technology to extract features and reduce the dimension of the vehicle-side raw data and construct a multidimensional state space including environmental parameters, system resources, and data attributes, the method further includes:

[0023] Obtaining the operating status data of multiple vehicle-side sensors respectively, and judging whether the corresponding vehicle-side sensors are operating normally based on the operating status data;

[0024] If it operates normally, data features are extracted from various vehicle-side raw data, and the data governance and processing logic of various vehicle-side raw data collected by various vehicle-side sensors is obtained respectively. The data correlation between the data collected by different vehicle-side sensors is determined. Based on the data correlation, all the collected data corresponding to different vehicle-side sensors are used as the basis to repeatedly group the collected vehicle-side raw data.

[0025] Based on the grouping results, the data features of the vehicle-side raw data in the same group are detected for correlation, and it is determined whether the data logic of the vehicle-side raw data collected by the target vehicle-side sensor and its corresponding vehicle-side raw data in the same group conforms to their corresponding data governance processing sub-logic;

[0026] If it meets the requirements, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal;

[0027] Otherwise, obtain the logical abnormality data corresponding to the suspected abnormal vehicle-side raw data of the target vehicle-side sensor and the target data logical detection result corresponding to the logical abnormality data. If the target data logical detection result is normal, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is abnormal, and the suspected abnormal vehicle-side raw data is marked as missing vehicle-side raw data;

[0028] If the target data logic detection result is abnormal, the data in the data group corresponding to the target vehicle-side sensor and the vehicle-side sensor corresponding to the logical abnormal data are compared to determine whether the two are exactly the same;

[0029] If they are exactly the same, it is determined that the suspected abnormal vehicle-side original data and its corresponding logical abnormal data are both missing vehicle-side original data;

[0030] If they are not exactly the same, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal;

[0031] If it is abnormal, the vehicle-mounted positioning information of the abnormal vehicle-side multi-sensor is obtained, and the missing vehicle-side original data is determined based on the vehicle-mounted positioning information.

[0032] Preferably, in a vehicle-side data management method based on reinforcement learning, the method further includes automatically supplementing missing vehicle-side original data, including:

[0033] Obtain historical data of missing vehicle-side raw data, sort the historical data based on the time axis to obtain a time data sequence, obtain the collection time characteristics of the current vehicle-side raw data, and filter the time data sequence based on the collection time characteristics to obtain a time data sequence;

[0034] Based on the data type and data presentation of the missing vehicle-side original data, combined with the data governance and processing logic corresponding to the missing vehicle-side original data, determine the current associated vehicle-side data corresponding to the missing vehicle-side original data;

[0035] respectively obtaining first similarities between a plurality of currently associated vehicle-side data and their corresponding historically associated vehicle-side data, and summing the first similarities to obtain a second similarity;

[0036] Compare and sort the second similarities of the historical associated vehicle-side data corresponding to each historical data in the time data sequence to obtain a historical data sequence, and use the historical data with the largest second similarity as reference data;

[0037] Based on the data governance processing logic, determine the data logic between each historical data and its corresponding historical associated vehicle-side data, as well as the data processing thread where the historical data and its corresponding historical associated vehicle-side data are located;

[0038] Align all historical data of the same data processing thread and its corresponding historical associated vehicle-side data, and compare them in combination with their corresponding data logic to determine the correlation change characteristics between the historical associated vehicle-side data of the same data processing thread and the corresponding historical data;

[0039] Obtain the correlation change features corresponding to multiple data processing threads and generate feature data clusters;

[0040] According to the characteristic data cluster, the reference data is corrected to obtain the supplementary data of the missing vehicle-side original data.

[0041] Preferably, in a vehicle-side data management method based on reinforcement learning, the reference data is corrected according to the characteristic data cluster to obtain supplementary data for the missing vehicle-side original data, including:

[0042] Obtaining differences between historical associated vehicle-side data of multiple data processing threads corresponding to the reference data and corresponding current associated vehicle-side data, and determining a difference amplitude characteristic of the associated vehicle-side data based on the differences;

[0043] Determining the error rate of the reference data based on the magnitude difference feature of the associated vehicle-side data and the difference in change features between the associated change features of the corresponding data processing threads within the feature data cluster;

[0044] When the error rates corresponding to all data processing threads are less than a preset threshold, the parameter data is determined to be supplementary data;

[0045] Otherwise, the error mean and standard deviation of the error rates corresponding to all data processing threads are calculated. When the standard deviation is less than a preset value, the reference data is corrected based on the error mean to obtain supplementary data.

[0046] When the standard deviation is greater than or equal to a preset value, an error extreme value is obtained, an error correction coefficient is obtained based on the error extreme value and the mean value, and the error mean value is corrected based on the error correction coefficient to obtain a corrected error rate;

[0047] The reference data is corrected based on the corrected error rate to obtain supplementary data.

[0048] Preferably, in a vehicle-side data governance method based on reinforcement learning, the current vehicle state is determined based on a multi-dimensional state space, the current state is analyzed through a multi-agent reinforcement learning model, and the target execution strategy corresponding to the current vehicle is determined to generate a target action, including:

[0049] Determine the current vehicle state of the vehicle based on the mapping positions of various vehicle raw data in multiple state spaces;

[0050] Input the current vehicle state into the multi-agent reinforcement learning model for analysis, and determine the target execution strategy for each agent based on the data governance participation process corresponding to various vehicle-side data and the policy execution responsibilities corresponding to each agent;

[0051] Based on the target execution strategy, the vehicle operation process corresponding to the intelligent agent is determined, and the vehicle operation process is mapped into specific continuous target actions and sent to the vehicle control system to perform the corresponding operations.

[0052] Preferably, in a vehicle-side data governance method based on reinforcement learning, the effect of the target action is evaluated based on a reward function, and based on the evaluation result, a feedback signal is generated and input into a multi-agent reinforcement learning model to optimize the model parameters, including:

[0053] Based on the reward function and the data attributes corresponding to the original data on the vehicle side, the effect of the target action is evaluated to obtain the risk assessment result of sensitive data leakage;

[0054] Based on the reward function and the changes in computing resource data during the execution of the target action, the efficiency of the target action execution and resource utilization are evaluated to obtain the governance energy efficiency evaluation results;

[0055] Generate feedback signals based on the sensitive data leakage risk assessment results and the governance energy efficiency assessment results and input them into the multi-agent reinforcement learning model;

[0056] After receiving the feedback signal, the multi-agent reinforcement learning model compares the sensitive data leakage risk assessment results and the governance energy efficiency assessment results with the model's original optimal sensitive data leakage risk assessment results and the governance energy efficiency assessment results respectively;

[0057] Determine whether the target action effect is the optimal effect. If so, collect the target action and its corresponding vehicle-side raw data as the latest training set to update the multi-agent reinforcement learning model and complete model parameter optimization.

[0058] The present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the vehicle-side data governance method based on reinforcement learning are implemented.

[0059] The present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program;

[0060] Among them, when the computer program is executed by the processor, the steps of the vehicle-side data governance method based on reinforcement learning are implemented.

[0061] Compared with the prior art, the present invention has at least the following beneficial effects:

[0062] The present invention collects real-time traffic environment data and vehicle status data of vehicles through multiple sensors on the vehicle side, which can comprehensively and accurately obtain various information during the operation of the vehicle, and integrate the traffic environment (such as congestion level, weather, traffic flow, road type, etc.), system resources (such as computing power utilization, storage remaining space, CPU / GPU remaining computing power, etc.), data attributes (such as privacy level, timeliness, data volume, data accuracy, etc.) cross-domain dynamic information in real time to form a complete portrayal of the data governance scenario, avoid computing power overload, and effectively reduce data processing delay and improve storage resource utilization compared to traditional fixed threshold strategies. A multi-agent enhanced learning model is established, so that each agent can perceive the global resource status of the vehicle through the multi-dimensional state space, and dynamically adjust the data collection frequency and processing priority. The system can autonomously discover the optimal solution for resource allocation and dynamically evaluate the data privacy level through artificial intelligence technology. The multi-agent reinforcement learning model can automatically match the encryption algorithm and storage area, achieve a balance between privacy protection and data utilization, and greatly reduce the risk of sensitive data leakage. It can also evaluate the effect of the target action based on the reward function, and generate a feedback signal based on the evaluation result to input the multi-agent reinforcement learning model to optimize the model parameters, so that the model can continuously learn and improve to adapt to different environments and task requirements. The reward function can be designed according to specific application scenarios and goals, such as improving driving safety, optimizing energy consumption, and improving traffic efficiency, thereby guiding the vehicle's decision-making in the desired direction and continuously improving the vehicle's performance and performance. The present invention covers the full life cycle management of data collection, processing, storage, and transmission, breaking through the limitations of static rules, and using artificial intelligence technology to achieve intelligent and adaptive data governance. By modeling the data governance process through the Markov model, the system can autonomously adapt to changes in the traffic environment, fluctuations in computing load, and policy compliance requirements, reducing the cost of manual intervention and improving the intelligence level of vehicle-side data governance.

[0063] The electronic device described in this invention significantly enhances the vehicle's intelligence by implementing a vehicle-side data management method based on reinforcement learning. This enables intelligent management of vehicle-side data, encompassing the entire process of data collection, processing, storage, and transmission. This effectively enhances the vehicle's autonomous perception, analysis, and decision-making capabilities, providing users with a more convenient, safe, and efficient travel experience. A computer-readable storage medium is also provided to provide a stable and reliable storage medium for the computer program implementing the vehicle-side data management method based on reinforcement learning.

[0064] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.

[0065] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0067] Figure 1 This is a flow chart of a vehicle-side data management method based on reinforcement learning of the present invention;

[0068] Figure 2 A flowchart for building a multi-agent reinforcement learning model;

[0069] Figure 3 Flowchart constructed for multidimensional state spaces. DETAILED DESCRIPTION

[0070] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0071] Example 1:

[0072] The present invention provides a vehicle-side data management method based on reinforcement learning, such as Figure 1 Shown, including:

[0073] Step 1: Collect the vehicle's real-time traffic environment data and vehicle status data through multiple sensors on the vehicle side to obtain the vehicle side raw data;

[0074] Step 2: Use artificial intelligence technology to extract features and reduce the dimensionality of the vehicle-side raw data, and construct a multidimensional state space that includes environmental parameters, system resources, and data attributes;

[0075] Step 3: Based on the multi-dimensional state space, determine the current vehicle state, analyze the current state through the multi-agent reinforcement learning model, determine the target execution strategy corresponding to the current vehicle and generate the target action;

[0076] Step 4: Evaluate the effect of the target action based on the reward function. Based on the evaluation results, generate a feedback signal and input it into the multi-agent reinforcement learning model to optimize the model parameters.

[0077] In this embodiment, the strategy includes data screening rules (such as dynamically adjusting the sensor acquisition frequency according to real-time road conditions), compression algorithm selection (such as matching the optimal compression scheme for data of different precisions), storage partition allocation (such as isolating sensitive data and storing it in an encrypted area) and transmission strategy optimization (such as selecting edge processing or cloud upload based on network status).

[0078] In this embodiment, the target actions include data screening, storage partition allocation, and optimized transmission strategy.

[0079] The beneficial effects of the above technical solution: The present invention collects real-time traffic environment data and vehicle status data of the vehicle through multiple sensors on the vehicle side, which can comprehensively and accurately obtain various information during the operation of the vehicle, and integrate the traffic environment (such as congestion level, weather, traffic flow, road type, etc.), system resources (such as computing power utilization, storage remaining space, CPU / GPU remaining computing power, etc.), data attributes (such as privacy level, timeliness, data volume, data accuracy, etc.) cross-domain dynamic information in real time to form a complete portrayal of the data governance scenario, avoid computing power overload, and effectively reduce data processing delays and improve storage resource utilization compared to traditional fixed threshold strategies. A multi-agent enhanced learning model is established, so that each agent can perceive the global resource status of the vehicle through a multi-dimensional state space and dynamically adjust data collection. Frequency, processing priority and storage path, the system can autonomously discover the optimal solution for resource allocation, and dynamically evaluate the data privacy level through artificial intelligence technology. The multi-agent reinforcement learning model can automatically match the encryption algorithm and storage area, achieve a balance between privacy protection and data utilization, and greatly reduce the risk of sensitive data leakage. It also evaluates the effect of the target action based on the reward function, and generates a feedback signal based on the evaluation result to input the multi-agent reinforcement learning model to optimize the model parameters, so that the model can continuously learn and improve to adapt to different environments and task requirements. The reward function can be designed according to specific application scenarios and goals, such as improving driving safety, optimizing energy consumption, and improving traffic efficiency, thereby guiding the vehicle's decision-making in the desired direction and continuously improving the vehicle's performance and performance. The present invention covers the full life cycle management of data collection, processing, storage, and transmission, breaking through the limitations of static rules, and using artificial intelligence technology to achieve intelligent and adaptive data governance. By modeling the data governance process through the Markov model, the system can autonomously adapt to changes in the traffic environment, fluctuations in computing load and policy compliance requirements, reducing the cost of manual intervention and improving the intelligence level of vehicle-side data governance.

[0080] Example 2:

[0081] Based on Example 1, a vehicle-side data management method based on reinforcement learning is proposed. Figure 2 As shown, it also includes:

[0082] An asynchronous advantage action evaluation algorithm is used to build a multi-agent reinforcement learning model, including:

[0083] Step A: Define each sensor node, data processing module and storage unit as an intelligent agent;

[0084] Step B: Based on the Markov decision process, establish the data governance task model corresponding to each agent and initialize it;

[0085] Step C: Control multiple agents to interact with the environment on different threads or processors, and collect multi-dimensional state space data, execution actions, and reward data of each agent during the environment interaction process;

[0086] Step D: Based on the update rules of the asynchronous advantage action evaluation algorithm, the data governance task model corresponding to each agent is updated using the multi-dimensional state space data, the executed actions, and the reward data for the executed actions;

[0087] Step E: Repeat steps CD until the accumulated reward corresponding to the strategy corresponding to the agent reaches the maximum value, and end the training of the data governance task model corresponding to the agent.

[0088] In this embodiment, environmental interaction refers to the interaction between each intelligent agent and the environment (such as real-time traffic data, system resource status, data processing load, etc.) according to the decision logic of its corresponding data governance task model.

[0089] In this embodiment, the reward data includes but is not limited to improved processing efficiency and reduced latency.

[0090] In this embodiment, the determination of the agent can be expanded according to user needs.

[0091] The beneficial effects of the above technical solution: The present invention defines sensor nodes, data processing modules and storage units as intelligent agents, realizes the functional decoupling of system components, reduces the complexity of the system, and then establishes data governance task models corresponding to different intelligent agents based on the Markov decision process, which provides a basis for the intelligent agents to continuously adjust their own strategies according to the feedback (state, reward, etc.) of the environment, while also making the system have certain fault tolerance and robustness. If an intelligent agent fails or malfunctions, other intelligent agents can still continue to work, ensuring that the basic functions of the system are not affected, and allowing multiple intelligent agents to interact with the environment in parallel on different threads or processors. Each intelligent agent explores the environment and collects data at the same time, so that the overall training process can be completed in a shorter time, effectively improving the training efficiency of the system, and based on the update rules of the asynchronous advantage action evaluation algorithm, the data governance task model corresponding to each intelligent agent is updated using multi-dimensional state space data, execution action and reward data for execution action, realizing continuous iterative update of the model, which helps to achieve the global optimality of the system.

[0092] Example 3:

[0093] Based on Example 1, step 1: use artificial intelligence technology to extract features and reduce the dimension of the vehicle-side raw data, and construct a multidimensional state space containing environmental parameters, system resources, and data attributes, such as Figure 3 Shown, including:

[0094] Step 101: Identify the vehicle-side raw data based on artificial intelligence technology to determine the data type corresponding to the vehicle-side raw data;

[0095] Step 102: Based on the data type, a neural network is selected using artificial intelligence technology to obtain the optimal dimensionality reduction network corresponding to the vehicle-side original data;

[0096] Step 103: performing data feature extraction and dimensionality reduction on the corresponding original vehicle-side data based on the optimal neural extraction network to obtain low-dimensional features of the original vehicle-side data;

[0097] Step 104: directly concatenate the low-dimensional features to obtain a state vector, and perform spatial mapping processing on the state vector to generate a multi-dimensional state space.

[0098] The beneficial effects of the above technical solution are as follows: the present invention uses artificial intelligence technology to identify the original data on the vehicle side, can accurately determine the category to which the data belongs, can clearly define data of different properties, provide a basis for subsequent targeted processing, and help improve the accuracy and efficiency of data processing. The optimal dimensionality reduction network is selected based on the data type, and different types of data use appropriate neural networks, which can give full play to the advantages of each network and make the dimensionality reduction process more in line with the characteristics of the data, thereby improving the dimensionality reduction effect and reducing information loss. The original vehicle-side data is extracted and reduced in dimension using the optimal neural extraction network, and the low-dimensional features are directly spliced into state vectors, and spatial mapping processing is performed to generate a multi-dimensional state space. While reducing the data dimension, the key information and internal structure of the original data are retained to the greatest extent. The multi-dimensional state space can more comprehensively and accurately describe the vehicle status, provide richer and more effective data support for the vehicle's intelligent decision-making, fault diagnosis, performance optimization, etc., and can better meet the data requirements of vehicle control systems, intelligent driving algorithms, etc., help to improve the vehicle's perception and response speed to complex environments, optimize the vehicle's control strategy and decision-making process, and thus improve the performance and safety of the entire vehicle system.

[0099] Example 4:

[0100] Based on Example 3, artificial intelligence technology is used to extract features and reduce the dimension of the vehicle-side raw data, and before constructing a multidimensional state space containing environmental parameters, system resources, and data attributes, the following steps are also included:

[0101] Obtaining the operating status data of multiple vehicle-side sensors respectively, and judging whether the corresponding vehicle-side sensors are operating normally based on the operating status data;

[0102] If it operates normally, data features are extracted from various vehicle-side raw data, and the data governance and processing logic of various vehicle-side raw data collected by various vehicle-side sensors is obtained respectively. The data correlation between the data collected by different vehicle-side sensors is determined. Based on the data correlation, all the collected data corresponding to different vehicle-side sensors are used as the basis to repeatedly group the collected vehicle-side raw data.

[0103] Based on the grouping results, the data features of the vehicle-side raw data in the same group are detected for correlation, and it is determined whether the data logic of the vehicle-side raw data collected by the target vehicle-side sensor and its corresponding vehicle-side raw data in the same group conforms to their corresponding data governance processing sub-logic;

[0104] If it meets the requirements, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal;

[0105] Otherwise, obtain the logical abnormality data corresponding to the suspected abnormal vehicle-side raw data of the target vehicle-side sensor and the target data logical detection result corresponding to the logical abnormality data. If the target data logical detection result is normal, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is abnormal, and the suspected abnormal vehicle-side raw data is marked as missing vehicle-side raw data;

[0106] If the target data logic detection result is abnormal, the data in the data group corresponding to the target vehicle-side sensor and the vehicle-side sensor corresponding to the logical abnormal data are compared to determine whether the two are exactly the same;

[0107] If they are exactly the same, it is determined that the suspected abnormal vehicle-side original data and its corresponding logical abnormal data are both missing vehicle-side original data;

[0108] If they are not exactly the same, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal;

[0109] If it is abnormal, the vehicle-mounted positioning information of the abnormal vehicle-side multi-sensor is obtained, and the missing vehicle-side original data is determined based on the vehicle-mounted positioning information.

[0110] In this embodiment, the data governance processing logic refers to the processing logic relationship of various data during the processing of the original data on the vehicle side, including the operation logic between the data and the data processing thread relationship.

[0111] In this embodiment, repeatable grouping means that any type of vehicle-side original data can be in multiple groups.

[0112] In this embodiment, the data governance processing sub-logic refers to the data instruction processing logic corresponding to the original data of the same group of vehicles.

[0113] In this embodiment, the suspected abnormal vehicle-side original data refers to the vehicle-side original data in the same group that does not conform to its corresponding data instruction processing sub-logic.

[0114] The beneficial effects of the above technical solution are as follows: the present invention can timely discover the faults or abnormal conditions of the sensors themselves by obtaining the operating status data of multiple vehicle-side sensors and judging whether they are operating normally, thereby avoiding data errors or missing due to sensor failures, thereby ensuring the reliability of the data collection link in the vehicle system and providing a reliable basis for subsequent data governance decisions; feature extraction is performed on multiple vehicle-side raw data, the data correlation between the data collected by different vehicle-side sensors is determined, and repeatable grouping is performed based on the correlation, which helps to deeply explore the intrinsic connection between the data of various vehicle-side sensors. A clear grasp of the logical relationship of the data can provide a more comprehensive understanding of the vehicle operating status, and provide richer and more accurate information for subsequent data analysis and decision-making; by performing correlation detection on the data features of the vehicle-side raw data in the same group, it is determined whether the data logic conforms to the corresponding data governance processing logic, and data anomalies can be accurately identified. For suspected abnormal data, detailed judgment is further performed through target data logic detection results and comparison with other sensor data groups to distinguish whether it is data missing or other abnormal conditions, thereby improving the accuracy and reliability of data anomaly judgment and avoiding misjudgment and omission. When it is found that the sensor is operating abnormally, the on-board positioning information of the abnormal vehicle-side multi-sensor is obtained to determine the missing vehicle-side original data. Combined with the vehicle's location information, the missing data can be found more specifically, which is conducive to subsequent data completion and repair, and helps to improve the integrity and availability of the data. The present invention can effectively discover and eliminate abnormal data, mark missing data, and significantly improve the quality of vehicle-side original data, provide strong support for the stable operation and intelligent development of the vehicle system, and provide a more accurate basis for the vehicle's intelligent decision-making. For example, in intelligent driving scenarios, accurate sensor data and reasonable data processing can enable the vehicle to perceive the surrounding environment more accurately, make more reasonable driving decisions, and improve the safety and efficiency of intelligent driving.

[0115] Example 5:

[0116] Based on Example 4, the automatic supplementation of missing vehicle-side original data includes:

[0117] Obtain historical data of missing vehicle-side raw data, sort the historical data based on the time axis to obtain a time data sequence, obtain the collection time characteristics of the current vehicle-side raw data, and filter the time data sequence based on the collection time characteristics to obtain a time data sequence;

[0118] Based on the data type and data presentation of the missing vehicle-side original data, combined with the data governance and processing logic corresponding to the missing vehicle-side original data, determine the current associated vehicle-side data corresponding to the missing vehicle-side original data;

[0119] respectively obtaining first similarities between a plurality of currently associated vehicle-side data and their corresponding historically associated vehicle-side data, and summing the first similarities to obtain a second similarity;

[0120] Compare and sort the second similarities of the historical associated vehicle-side data corresponding to each historical data in the time data sequence to obtain a historical data sequence, and use the historical data with the largest second similarity as reference data;

[0121] Based on the data governance processing logic, determine the data logic between each historical data and its corresponding historical associated vehicle-side data, as well as the data processing thread where the historical data and its corresponding historical associated vehicle-side data are located;

[0122] Align all historical data of the same data processing thread and its corresponding historical associated vehicle-side data, and compare them in combination with their corresponding data logic to determine the correlation change characteristics between the historical associated vehicle-side data of the same data processing thread and the corresponding historical data;

[0123] Obtain the correlation change features corresponding to multiple data processing threads and generate feature data clusters;

[0124] According to the characteristic data cluster, the reference data is corrected to obtain the supplementary data of the missing vehicle-side original data.

[0125] In this embodiment, the collection time feature refers to the collection period characteristics of the missing vehicle-side original data, for example, the morning peak period (7:30-9:00), the evening peak period (17:30-19:30), etc.

[0126] In this embodiment, the time data sequence refers to a sequence generated by arranging historical data with missing vehicle-side original data collection time features in chronological order.

[0127] In this embodiment, data representation can be divided into structured data, unstructured element data, numerical data, image data, etc.

[0128] In this embodiment, the first similarity refers to the similarity between the current associated vehicle-side data and the historical vehicle-side associated data with the same data processing thread (ie, process).

[0129] In this embodiment, the preset value interval is set to [0.8, 1].

[0130] In this embodiment, the data value interval refers to the interval range consisting of the historical data value corresponding to the historical associated vehicle-side data with the second highest similarity to the historical data value corresponding to the historical associated vehicle-side data with the second lowest similarity.

[0131] In this embodiment, data logic refers to the data governance processing logic between the missing vehicle-side original data and its corresponding associated vehicle-side data.

[0132] In this embodiment, the associated change feature refers to the corresponding change pattern of the historical data when the historical associated vehicle-side data in the same data processing thread as the historical data changes. For example, when the historical data is numerical data, if the historical associated vehicle-side data becomes larger, the corresponding historical data also becomes larger.

[0133] The beneficial effects of the above technical solution: The present invention obtains historical data of missing vehicle-side original data, combines timeline sorting and screening, and can mine information related to missing data from historical data, providing a basis for data supplementation. Based on the data type, expression form and governance logic, the current associated vehicle-side data is determined, and the similarity is calculated to screen the historical data, and then the reference data is determined. In the process of determining the supplementary data, the data type, expression form, governance logic and the correlation change characteristics between the data are fully considered. At the same time, by calculating the similarity to screen the appropriate historical data, and aligning and comparing the data in the same data processing thread, the value of the missing data can be inferred more accurately, reducing the blindness and arbitrariness of data supplementation and improving the accuracy of the supplementary data. Accurate data is crucial for various functions of the vehicle (such as fault diagnosis, performance evaluation, etc.), and helps to improve the reliability of the vehicle system. In the process of processing missing data, the correlation relationship between the data was deeply analyzed, including the logic between historical data and related vehicle-side data and the correlation change characteristics within the data processing thread. The mining of potential relationships in the data not only helps to supplement the missing data, but also discovers the hidden information and patterns in the data, providing deeper support for the optimized operation of the vehicle, intelligent decision-making, etc., fully considering the diversity and complexity of vehicle-side data, and using multiple factors to determine the supplementary data for missing data. It can better adapt to the complex environment of vehicle-side data. Whether it is for structured data or unstructured data, it can provide an effective missing data processing solution, enhancing the versatility and adaptability of the method.

[0134] Example 6:

[0135] On the basis of Example 5, the reference data is corrected according to the characteristic data cluster to obtain supplementary data for the missing vehicle-side original data, including:

[0136] Obtaining differences between historical associated vehicle-side data of multiple data processing threads corresponding to the reference data and corresponding current associated vehicle-side data, and determining a difference amplitude characteristic of the associated vehicle-side data based on the differences;

[0137] Determining the error rate of the reference data based on the magnitude difference feature of the associated vehicle-side data and the difference in change features between the associated change features of the corresponding data processing threads within the feature data cluster;

[0138] When the error rates corresponding to all data processing threads are less than a preset threshold, the parameter data is determined to be supplementary data;

[0139] Otherwise, the error mean and standard deviation of the error rates corresponding to all data processing threads are calculated. When the standard deviation is less than a preset value, the reference data is corrected based on the error mean to obtain supplementary data.

[0140] When the standard deviation is greater than or equal to a preset value, an error extreme value is obtained, an error correction coefficient is obtained based on the error extreme value and the mean value, and the error mean value is corrected based on the error correction coefficient to obtain a corrected error rate;

[0141] The reference data is corrected based on the corrected error rate to obtain supplementary data.

[0142] In this embodiment, the difference amplitude characteristics of the associated vehicle-side data include the difference amplitude and the difference direction, wherein the difference direction includes positive difference and negative difference. When the representation content of the historical associated vehicle-side data is greater than its corresponding current associated vehicle-side data, it is a negative difference; otherwise, it is a positive difference.

[0143] In this embodiment, the extreme error values refer to the maximum and minimum error rates corresponding to all data processing threads.

[0144] In this embodiment, the error correction coefficient is a quotient obtained by summing the differences between the maximum error value and the minimum error value and the mean error value and then dividing the sum by the mean error value.

[0145] The beneficial effects of the above technical solution: In the process of determining the supplementary data, the present invention takes into account the difference between the original and currently associated vehicle-side data corresponding to the reference data, and dynamically evaluates the reliability of the reference data. When the error rate does not meet the preset conditions, the reference data is adjusted to remove data with large errors and low reliability, and retain or correct data with small errors as supplementary data, which helps to improve the overall quality of vehicle-side data and provide a reliable basis for the accuracy of data governance decisions. It not only considers the error mean, but also introduces the standard deviation to measure the degree of dispersion of the data. When the standard deviation is less than the preset value, correction is made based on the error mean; when the standard deviation is greater than or equal to the preset value, the error correction coefficient is obtained by calculating the error extreme value to correct the error mean, ensuring that the supplementary data is consistent with the existing data in characteristics and logic, making the entire vehicle-side data more coherent and unified between different processing threads, avoiding analysis deviations and decision-making errors caused by data inconsistency. The present invention further analyzes the fluctuation of the data by calculating the error mean, standard deviation, and error extreme value, and corrects the reference data accordingly, effectively adapting to the dynamic changes of vehicle-side data under different times and conditions, making the supplementary data more in line with actual conditions, enhancing the adaptability of the data processing method to complex and changeable data environments, and helping to improve the performance and decision accuracy of the intelligent decision-making system, thereby improving the safety, efficiency and intelligence level of the vehicle.

[0146] Example 7:

[0147] On the basis of Example 1, the current vehicle state is determined based on the multidimensional state space, and the current state is analyzed through the multi-agent reinforcement learning model to determine the target execution strategy corresponding to the current vehicle and generate the target action, including:

[0148] Determine the current vehicle state of the vehicle based on the mapping positions of various vehicle raw data in multiple state spaces;

[0149] Input the current vehicle state into the multi-agent reinforcement learning model for analysis, and determine the target execution strategy for each agent based on the data governance participation process corresponding to various vehicle-side data and the policy execution responsibilities corresponding to each agent;

[0150] Based on the target execution strategy, the vehicle operation process corresponding to the intelligent agent is determined, and the vehicle operation process is mapped into specific continuous target actions and sent to the vehicle control system to perform the corresponding operations.

[0151] In this embodiment, the policy execution responsibility refers to the tasks that each agent undertakes in the data quality process. For example, the storage unit needs to undertake the task of storing data and partitioning the storage of sensitive data.

[0152] In this embodiment, the data governance participation process data has the beneficial effect of the above technical solution in the data processing process of different vehicle-side data: the present invention determines the current vehicle state by analyzing the mapping position of various vehicle raw data in multiple state spaces, and can comprehensively consider various aspects of vehicle information (environmental parameters, system resources, data attributes), and comprehensively and accurately describe the real-time state of the vehicle from multiple dimensions. It provides a reliable basis for subsequent decision-making and operations, so that the vehicle control system can make reasonable responses based on more real and detailed vehicle state information. Afterwards, the current vehicle state is input into the multi-agent reinforcement learning model for analysis, and the target execution strategy is determined according to the data governance participation process and the agent strategy execution responsibility. This enables each agent to make decisions that are more in line with the overall system optimization goals based on its own responsibilities and the actual state of the vehicle. At the same time, the collaboration and division of labor between multiple agents are clearer, avoiding the blindness and conflict of decision-making, and improving the scientificity and effectiveness of decision-making. Moreover, the multi-agent reinforcement learning model can process the decision-making process of multiple agents in parallel, using different threads or processors to simultaneously interact with the environment and learn strategies. Compared with traditional single-agent or sequential decision-making methods, it greatly improves the decision-making speed and can respond to vehicle state changes in a shorter time, meeting the needs of real-time vehicle control. It is particularly suitable for complex and changing traffic environments. Then, based on the data governance participation process and agent strategy execution responsibility, the target execution strategy is determined, enabling the system to flexibly adjust the decision-making methods and strategies of the agents according to different application scenarios and task requirements. When the vehicle faces different driving conditions (such as road conditions and traffic rules) or performs different tasks (such as cargo transportation and passenger pick-up), the system can quickly adapt and make corresponding decision adjustments, improving the versatility and adaptability of the system. The target execution strategy is mapped into specific continuous target actions and sent to the vehicle control system, ensuring the consistency and accuracy of vehicle operation, avoiding vehicle operation instability or safety hazards caused by discontinuous or unreasonable operation instructions. At the same time, based on accurate vehicle status and optimized decision-making, the vehicle operation can be more precise, improving driving comfort and safety, such as smoother acceleration, deceleration, and steering operations.

[0153] Example 8:

[0154] Based on Example 1, the target action effect is evaluated based on the reward function. Based on the evaluation result, a feedback signal is generated and input into the multi-agent reinforcement learning model to optimize the model parameters, including:

[0155] Based on the reward function and the data attributes corresponding to the original data on the vehicle side, the effect of the target action is evaluated to obtain the risk assessment result of sensitive data leakage;

[0156] Based on the reward function and the changes in computing resource data during the execution of the target action, the efficiency of the target action execution and resource utilization are evaluated to obtain the governance energy efficiency evaluation results;

[0157] Generate feedback signals based on the sensitive data leakage risk assessment results and the governance energy efficiency assessment results and input them into the multi-agent reinforcement learning model;

[0158] After receiving the feedback signal, the multi-agent reinforcement learning model compares the sensitive data leakage risk assessment results and the governance energy efficiency assessment results with the model's original optimal sensitive data leakage risk assessment results and the governance energy efficiency assessment results respectively;

[0159] Determine whether the target action effect is the optimal effect. If so, collect the target action and its corresponding vehicle-side raw data as the latest training set to update the multi-agent reinforcement learning model and complete model parameter optimization.

[0160] The beneficial effects of the above technical solution: The present invention evaluates the risk of sensitive data leakage by combining the data attributes and reward functions corresponding to the original data on the vehicle side, which can fully consider the characteristics of the data (such as the sensitivity and importance of the data), comprehensively and accurately measure the security risks that may be brought about by the target action, evaluate the protection of user privacy data by the current model decision and incorporate it into the feedback mechanism, prompting the multi-agent reinforcement learning model to pay more attention to data security in subsequent decision-making, which is conducive to achieving a balance between privacy protection and data utilization, and greatly reducing the risk of sensitive data leakage; based on the reward function and combined with the changes in computing power resource data, the target action execution efficiency and resource utilization are evaluated, and the action effect can be quantitatively analyzed from the perspective of resource consumption and execution efficiency, which helps to discover problems of resource waste or inefficiency and achieve reasonable resource allocation. The two evaluation results generate feedback signals and input them into the multi-agent reinforcement learning model, and compared with the original optimal results, so that the model can continuously reflect on and improve its own decisions. The model can judge whether the current action is more effective based on the comparison results, and thus gradually adjust the strategy, tending to choose actions that can both reduce the risk of data leakage and improve governance efficiency, thereby improving the scientificity and rationality of the model's decision-making. When the target action is judged to be optimal, the target action and its corresponding vehicle-side original data are collected as the latest training set to update the model and train it, thereby optimizing the model parameters. This allows the model to continuously adapt to new situations and data changes, continuously improve its performance and accuracy, and provide a basis for the model to better cope with the complex and changing vehicle operating environment and data governance needs.

[0161] Example 9:

[0162] The present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor:

[0163] When the processor executes the computer program, the steps of the vehicle-side data governance method based on reinforcement learning are implemented.

[0164] The beneficial effects of the above technical solution: The electronic device described in the present invention can significantly improve the intelligence level of the vehicle by implementing a vehicle-side data management method based on reinforcement learning, and realize intelligent management of vehicle-side data covering the entire process of data collection, processing, storage, and transmission, effectively promoting the vehicle's ability to have autonomous perception, analysis and decision-making, and providing users with a more convenient, safe and efficient travel experience.

[0165] Example 10:

[0166] The present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program;

[0167] Among them, when the computer program is executed by the processor, the steps of the vehicle-side data governance method based on reinforcement learning are implemented.

[0168] The beneficial effects of the above technical solution: The present invention provides a computer-readable storage medium that provides a stable and reliable storage carrier for the computer program of the vehicle-side data governance method based on reinforcement learning.

[0169] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A vehicle-side data management method based on reinforcement learning, characterized in that: include: The vehicle's real-time traffic environment data, vehicle status data, and computing resource data are collected through multiple sensors on the vehicle to obtain the vehicle's original data; Use artificial intelligence technology to extract features and reduce the dimension of vehicle-side raw data, and construct a multi-dimensional state space that includes environmental parameters, system resources, and data attributes; Based on the multi-dimensional state space, the current vehicle state is determined, and the current state is analyzed through the multi-agent reinforcement learning model to determine the target execution strategy corresponding to the current vehicle and generate the target action; The effect of the target action is evaluated based on the reward function. Based on the evaluation results, a feedback signal is generated and input into the multi-agent reinforcement learning model to optimize the model parameters.

2. The vehicle-side data management method based on reinforcement learning according to claim 1 is characterized in that: Also includes: An asynchronous advantage action evaluation algorithm is used to build a multi-agent reinforcement learning model, including: Step A: Define each sensor node, data processing module and storage unit as an intelligent agent; Step B: Based on the Markov decision process, establish the data governance task model corresponding to each agent and initialize it; Step C: Control multiple agents to interact with the environment on different threads or processors, and collect multi-dimensional state space data, execution actions, and reward data of each agent during the environment interaction process; Step D: Based on the update rules of the asynchronous advantage action evaluation algorithm, the data governance task model corresponding to each agent is updated using the multi-dimensional state space data, the executed actions, and the reward data for the executed actions; Step E: Repeat steps CD until the accumulated reward corresponding to the strategy corresponding to the agent reaches the maximum value, and end the training of the data governance task model corresponding to the agent.

3. The vehicle-side data management method based on reinforcement learning according to claim 1 is characterized in that: Artificial intelligence technology is used to extract features and reduce the dimensionality of raw vehicle data, constructing a multidimensional state space that includes environmental parameters, system resources, and data attributes, including: Identify the vehicle-side raw data based on artificial intelligence technology and determine the data type corresponding to the vehicle-side raw data; Based on the data type, artificial intelligence technology is used to select a neural network to obtain the optimal dimensionality reduction network corresponding to the original data on the vehicle side; Based on the optimal neural extraction network, data feature extraction and dimensionality reduction are performed on the corresponding original vehicle-side data to obtain the low-dimensional features of the original vehicle-side data; The low-dimensional features are directly spliced to obtain a state vector, and the state vector is spatially mapped to generate a multi-dimensional state space.

4. The vehicle-side data management method based on reinforcement learning according to claim 3 is characterized in that: Before using artificial intelligence technology to extract features and reduce the dimensionality of raw vehicle data and construct a multidimensional state space that includes environmental parameters, system resources, and data attributes, the following steps are also required: Obtaining the operating status data of multiple vehicle-side sensors respectively, and judging whether the corresponding vehicle-side sensors are operating normally based on the operating status data; If it operates normally, data features are extracted from various vehicle-side raw data, and the data governance and processing logic of various vehicle-side raw data collected by various vehicle-side sensors is obtained respectively. The data correlation between the data collected by different vehicle-side sensors is determined. Based on the data correlation, all the collected data corresponding to different vehicle-side sensors are used as the basis to repeatedly group the collected vehicle-side raw data. Based on the grouping results, the data features of the vehicle-side raw data in the same group are detected for correlation, and it is determined whether the data logic of the vehicle-side raw data collected by the target vehicle-side sensor and its corresponding vehicle-side raw data in the same group conforms to their corresponding data governance processing sub-logic; If it meets the requirements, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal; Otherwise, obtain the logical abnormality data corresponding to the suspected abnormal vehicle-side raw data of the target vehicle-side sensor and the target data logical detection result corresponding to the logical abnormality data. If the target data logical detection result is normal, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is abnormal, and the suspected abnormal vehicle-side raw data is marked as missing vehicle-side raw data; If the target data logic detection result is abnormal, the data in the data group corresponding to the target vehicle-side sensor and the vehicle-side sensor corresponding to the logical abnormal data are compared to determine whether the two are exactly the same; If they are exactly the same, it is determined that the suspected abnormal vehicle-side original data and its corresponding logical abnormal data are both missing vehicle-side original data; If they are not exactly the same, it is determined that the vehicle-side raw data collected by the target vehicle-side sensor is normal; If it is abnormal, the vehicle-mounted positioning information of the abnormal vehicle-side multi-sensor is obtained, and the missing vehicle-side original data is determined based on the vehicle-mounted positioning information.

5. The vehicle-side data management method based on reinforcement learning according to claim 4 is characterized in that: It also includes automatic supplementation of missing vehicle-side original data, including: Obtain historical data of missing vehicle-side raw data, sort the historical data based on the time axis to obtain a time data sequence, obtain the collection time characteristics of the current vehicle-side raw data, and filter the time data sequence based on the collection time characteristics to obtain a time data sequence; Based on the data type and data presentation of the missing vehicle-side original data, combined with the data governance and processing logic corresponding to the missing vehicle-side original data, determine the current associated vehicle-side data corresponding to the missing vehicle-side original data; respectively obtaining first similarities between a plurality of currently associated vehicle-side data and their corresponding historically associated vehicle-side data, and summing the first similarities to obtain a second similarity; Compare and sort the second similarities of the historical associated vehicle-side data corresponding to each historical data in the time data sequence to obtain a historical data sequence, and use the historical data with the largest second similarity as reference data; Based on the data governance processing logic, determine the data logic between each historical data and its corresponding historical associated vehicle-side data, as well as the data processing thread where the historical data and its corresponding historical associated vehicle-side data are located; Align all historical data of the same data processing thread and its corresponding historical associated vehicle-side data, and compare them in combination with their corresponding data logic to determine the correlation change characteristics between the historical associated vehicle-side data of the same data processing thread and the corresponding historical data; Obtain the correlation change features corresponding to multiple data processing threads and generate feature data clusters; According to the characteristic data cluster, the reference data is corrected to obtain the supplementary data of the missing vehicle-side original data.

6. The vehicle-side data management method based on reinforcement learning according to claim 5 is characterized in that: Based on the characteristic data cluster, the reference data is corrected to obtain the supplementary data of the missing vehicle-side original data, including: Obtaining differences between historical associated vehicle-side data of multiple data processing threads corresponding to the reference data and corresponding current associated vehicle-side data, and determining a difference amplitude characteristic of the associated vehicle-side data based on the differences; Determining the error rate of the reference data based on the magnitude difference feature of the associated vehicle-side data and the difference in change features between the associated change features of the corresponding data processing threads within the feature data cluster; When the error rates corresponding to all data processing threads are less than a preset threshold, the parameter data is determined to be supplementary data; Otherwise, the error mean and standard deviation of the error rates corresponding to all data processing threads are calculated. When the standard deviation is less than a preset value, the reference data is corrected based on the error mean to obtain supplementary data. When the standard deviation is greater than or equal to a preset value, an error extreme value is obtained, an error correction coefficient is obtained based on the error extreme value and the mean value, and the error mean value is corrected based on the error correction coefficient to obtain a corrected error rate; The reference data is corrected based on the corrected error rate to obtain supplementary data.

7. The vehicle-side data management method based on reinforcement learning according to claim 1 is characterized in that: Based on the multi-dimensional state space, the current vehicle state is determined, and the current state is analyzed through the multi-agent reinforcement learning model to determine the target execution strategy corresponding to the current vehicle and generate the target action, including: Determine the current vehicle state of the vehicle based on the mapping positions of various vehicle raw data in multiple state spaces; Input the current vehicle state into the multi-agent reinforcement learning model for analysis, and determine the target execution strategy for each agent based on the data governance participation process corresponding to various vehicle-side data and the policy execution responsibilities corresponding to each agent; Based on the target execution strategy, the vehicle operation process corresponding to the intelligent agent is determined, and the vehicle operation process is mapped into specific continuous target actions and sent to the vehicle control system to perform the corresponding operations.

8. The vehicle-side data management method based on reinforcement learning according to claim 1 is characterized in that: The target action effect is evaluated based on the reward function. Based on the evaluation results, a feedback signal is generated and input into the multi-agent reinforcement learning model to optimize the model parameters, including: Based on the reward function and the data attributes corresponding to the original data on the vehicle side, the effect of the target action is evaluated to obtain the risk assessment result of sensitive data leakage; Based on the reward function and the changes in computing resource data during the execution of the target action, the efficiency of the target action execution and resource utilization are evaluated to obtain the governance energy efficiency evaluation results; Generate feedback signals based on the sensitive data leakage risk assessment results and the governance energy efficiency assessment results and input them into the multi-agent reinforcement learning model; After receiving the feedback signal, the multi-agent reinforcement learning model compares the sensitive data leakage risk assessment results and the governance energy efficiency assessment results with the model's original optimal sensitive data leakage risk assessment results and the governance energy efficiency assessment results respectively; Determine whether the target action effect is the optimal effect. If so, collect the target action and its corresponding vehicle-side raw data as the latest training set to update the multi-agent reinforcement learning model and complete model parameter optimization.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the vehicle-side data governance method based on reinforcement learning described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program; Wherein, when the computer program is executed by the processor, the steps of the vehicle-side data governance method based on reinforcement learning described in any one of claims 1 to 8 are implemented.