Intelligent driving cooperative control method and system based on cloud edge cooperation

Through the Attention-GAN model and lightweight execution network that collaborates with cloud edge, the data transmission and computing resource limitation of the autonomous driving system is solved, efficient intelligent driving control is achieved, and the system's safety and adaptability is improved.

CN120348312APending Publication Date: 2025-07-22CHANGCHUN AUTOMOTIVE TEST CENT
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510431987.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing autonomous driving system consumes time and bandwidth in data upload, expensive cloud computing resources, limited computing power of on-board equipment, and personal privacy data security issues, making it difficult to achieve efficient cloud-side collaborative intelligent driving control.

Method used

Design an intelligent driving collaborative control method based on cloud-edge collaboration, use the Attention-GAN model to jointly learn perception, decision-making and control strategies in the cloud, and deploy a lightweight execution network on the vehicle end, optimize data transmission through incremental updates and selective communications, and realize distributed collaborative optimization in combination with federated learning.

Benefits of technology

It improves the safety, robustness and generalization performance of the autonomous driving system, reduces communication overhead, improves the real-time and scalability of the system, and adapts to complex and changeable traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120348312A_ABST
    Figure CN120348312A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent driving cooperative control method and system based on cloud side cooperation, and the method and system achieve the behavior decision optimization of cloud side-vehicle side cooperation through Attention-GAN. According to the framework, firstly, an attention mechanism is utilized to enable a model to adaptively pay attention to key information in an environment, and a GAN is introduced to optimize a driving strategy, so that the accuracy and robustness of decision making are improved. On the basis, a cloud-edge heterogeneous collaborative system architecture is established, distributed collaborative optimization is realized through federal learning, the powerful computing power of the cloud end can be utilized to learn general rules from mass data, and the real-time processing capacity of the vehicle end can be utilized to make personalized decisions. And meanwhile, strategies such as model increment updating and selective communication uploading are designed, so that the privacy of the user is protected while the communication overhead is saved. According to the method, the safety, robustness and generalization performance of the automatic driving system can be remarkably improved, and the wide application prospect of combination of deep learning and cloud edge calculation in the field of automatic driving is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving control, and particularly to an intelligent driving collaborative control method and system based on cloud-edge collaboration. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, autonomous driving has become a research hotspot in the transportation field. Intelligent connected vehicles can significantly improve traffic safety, traffic efficiency and comfort, and are of great significance for alleviating traffic congestion, reducing accidents and saving drivers' time. In an intelligent driving system, perception, decision-making and control are its key technologies. Among them, the perception technology is responsible for obtaining environmental information around the vehicle, such as the position of obstacles, lane lines, traffic signs, etc.; the decision-making technology conducts path planning and behavior decision-making based on the perception information; and the control technology is responsible for executing the decision-making instructions and precisely controlling the throttle, brakes and steering of the vehicle.

[0003] Traditional autonomous driving perception algorithms are mainly based on computer vision and sensor fusion technologies, such as detecting lane lines by Hough transform, detecting traffic signs by Haar features, and fusing lidar and visual information by Kalman filtering, etc. These algorithms have good effects in structured road environments, but are less robust in unstructured environments such as complex urban roads and bad weather. Deep learning methods can automatically extract multi-level features from massive data, overcome the dependence of traditional algorithms on prior knowledge, and have made breakthrough progress in visual tasks such as image classification, object detection and semantic segmentation.

[0004] Autonomous driving perception technology based on deep learning has gradually become the mainstream. Companies such as Tesla have taken the lead in proposing an end-to-end driving model, directly predicting the steering wheel angle from in-vehicle camera images and realizing the road keeping function. Companies such as Waymo and Baidu have also carried out large-scale road tests to verify the effectiveness of deep learning in autonomous driving.

[0005] However, most current deep learning perception models adopt a centralized learning paradigm, that is, uploading the data collected by in-vehicle devices to the cloud for training, and then deploying the trained model to the vehicle end for inference. This mode has the following problems:

[0006] (1) Data uploading is time-consuming and bandwidth-consuming, and it is difficult to process massive in-vehicle data in real time;

[0007] (2) Cloud computing resources are expensive and the training cost is high;

[0008] (3) The computing power of in-vehicle devices is limited and it is difficult to run complex deep models;

[0009] (4) Uploading personal privacy data to the cloud has security risks.

[0010] Therefore, there is an urgent need for an intelligent driving control framework for cloud-edge collaboration, which can utilize powerful computing resources in the cloud to train advanced perception, decision-making, and control models, and efficiently deploy the models to in-vehicle devices for inference. At the same time, the incremental data generated by the vehicle can be processed locally and used to continuously optimize the cloud model, rather than uploading all of it to the cloud. In addition, distributed deep learning for cloud-edge collaboration has received extensive attention in the academic and industrial fields in recent years. Google proposed the concept of Federated Learning in 2016, which aggregates local models on different devices through encrypted communication to achieve federated training of the model without revealing the original data. However, the application of federated learning for cloud-edge collaboration to autonomous driving is still in its infancy. Summary of the Invention

[0011] In view of the above problems, the present invention proposes an intelligent driving collaborative control method and system based on cloud-edge collaboration, and designs a GAN enhanced by the Attention mechanism to jointly learn the perception, decision-making, and control strategies for autonomous driving. The Generator is responsible for generating realistic driving behaviors, and the Discriminator is responsible for distinguishing real and generated behaviors. The two play against each other to make the strategy more robust. The Attention mechanism can model the dependence relationship between driving behaviors and key traffic objects to generate safer and more efficient decisions. At the vehicle end, a lightweight execution network is designed to run the cloud strategy model and control the vehicle. At the same time, the execution network can collect in-vehicle sensor data, extract key features and upload them to the cloud for continuous optimization of the strategy model.

[0012] The technical solution of the present invention is as follows: An intelligent driving collaborative control method based on cloud-edge collaboration, comprising the following steps:

[0013] Step 1: The cloud-edge system collects and preprocesses data related to intelligent driving, specifically including:

[0014] Step 1.1: Collect autonomous driving data in the real road environment, including time series information such as in-vehicle camera images, vehicle speed, and acceleration.

[0015] Step 1.2: Perform preprocessing on the original images, including enhancement, denoising, and correction, and unify the size to h×w×c, where h represents the image height, w represents the image width, and c represents the number of image channels.

[0016] Step 1.3: Divide the data set into a training set, a validation set, and a test set according to a ratio of 8:1:1.

[0017] Step 2: Build a cloud Attention-GAN strategy model, specifically including:

[0018] Step 2.1: Construct Generator G, which generates a false driving behavior based on the current environmental state s cnn , cnn , t , t , t , t , t , t and the historical behavior a t-1 where

[0019]

[0020] denotes the generated driving behavior, s represents the current environmental state, a t represents the historical behavior, and θ t-1 is the parameter of Generator; g

[0021] Step 2.2: Introduce a visual feature extractor f cnn (·) to extract the compact features of the image:

[0022]

[0023] where F t denotes the extracted image features, f cnn (·) represents the convolutional neural network feature extractor function, I t is the input original image, h′ represents the height of the feature map, w′ represents the width of the feature map, and c′ represents the number of channels of the feature map.

[0024] Step 2.3: Design an attention mechanism to calculate weights through query, key, and value, and obtain attention features by weighting the feature map

[0025]

[0026] where A t denotes the attention weight matrix, represents the weighted attention features, softmax(·) represents the softmax normalization function, Q t denotes the query matrix, K t denotes the key matrix, V t denotes the value matrix, and c″ is the scaling factor.

[0027] Step 2.4: Input the attention features and the state information into an LSTM to model the temporal dependence and generate the final driving behavior

[0028] Step 2.5: Construct a Discriminator to distinguish between real and generated behavior sequences:

[0029]

[0030] Among them, represents the authenticity score, and D(·) represents the Discriminator function. represents the time-series data of the environmental state and behavior, T represents the time-series length, and θ d is the parameter of the Discriminator.

[0031] Step 3: Cloud-based GEnerative-Adversarial training, specifically including:

[0032] Step 3.1: Train the Discriminator to maximize the log probability of real behaviors and the negative log probability of fake behaviors:

[0033]

[0034] Among them, represents the loss function of the Discriminator, m is the batch size, and s i represents the i-th real environmental state, and a i represents the i-th real behavior. represents the i-th fake behavior generated by the Generator.

[0035] Step 3.2: Train the Generator to minimize the negative log probability of fake behaviors:

[0036]

[0037] Among them, represents the loss function of the Generator.

[0038] Step 3.3: Alternately optimize the Generator and the Discriminator until the model converges.

[0039] Step 4: Perform network deployment on the vehicle side, including:

[0040] Step 4.1: Design a lightweight policy network π, which takes the environmental state s t , historical behavior a t-1 as inputs and outputs the optimal driving behavior distribution μ t :

[0041] μ t = π(s t , a t-1 ; θ π );

[0042] Among them, μ t represents the generated behavior distribution, π(·) represents the policy network function, and θπ are the parameters of the policy network. Sample a driving instruction a from μ t t .

[0043] Step 4.2: Optimize the policy network using model compression techniques to control the computational overhead.

[0044] Step 4.3: Develop a vehicle - side environment perception module to extract key driving information and upload it to the cloud.

[0045] Step 5: Design the cloud - edge communication mechanism, including:

[0046] Step 5.1: Synchronize the cloud - side and vehicle - side models in an incremental update manner, that is, only transmit the increment Δθ k :

[0047] Δθ k = θ k - θ k-1 ;

[0048]

[0049] where Δθ k represents the increment of the model parameters in the k - th update, θ k represents the cloud - side model parameters after the k - th update, θ k-1 represents the cloud - side model parameters after the (k - 1) - th update, represents the vehicle - side model parameters after the k - th update. And perform sparse quantization compression on the increment.

[0050] Step 5.2: The vehicle - side selectively uploads environmental features according to the communication bandwidth B and the feature importance metric ρ:

[0051]

[0052] where is the subset of the feature dimensions to be uploaded, ρ i represents the importance metric of the i - th dimension feature, γ∈(0, 1] is the upload ratio, and d represents the dimension of the feature.

[0053] Step 5.3: Design a federated learning framework, where the cloud aggregates the local updates of different vehicles to optimize the global policy:

[0054]

[0055] where f(θ) is the global objective function, F i (θ) is the local objective function of the i - th vehicle, N is the total number of vehicles participating in federated learning, represents the local dataset of the i - th vehicle, represents the dataset of all vehicles.​

[0056] In addition, the present invention also proposes an intelligent driving collaborative control system based on cloud-edge collaboration, which consists of the following key modules:

[0057] 1. Cloud-edge collaborative intelligent driving data acquisition and preprocessing module:

[0058] This module is used to collect multi-source heterogeneous data required for autonomous driving, including sequential information such as in-vehicle camera images, vehicle speed, acceleration, etc. And this module cleans, annotates, enhances, and preprocesses the original data, divides the data set into a training set, a validation set, and a test set for model learning and evaluation.

[0059] 2. Cloud-based policy learning module, including:

[0060] Generator based on attention mechanism: Generates safe and smooth driving behavior sequences according to the current environmental state and historical behaviors.

[0061] Discriminator based on convolutional neural network: Distinguishes real and generated behavior sequences to guide the optimization of the Generator.

[0062] Through Generatove-Adversarial training (GAN), the cloud continuously optimizes and improves the policy model of the Generator and learns general driving knowledge.

[0063] 3. Vehicle-side policy execution module, including:

[0064] Lightweight policy network: Compresses and deploys the policy model trained on the cloud to in-vehicle devices, receives the environmental state in real time, and outputs driving behavior decisions.

[0065] Environmental perception unit: Uses in-vehicle sensors to collect traffic environment information in real time, extracts and uploads key features to the cloud to optimize data transmission efficiency.

[0066] Policy execution unit: Executes the driving decisions generated by the policy network, controls actuators such as the throttle, brakes, and steering, and realizes the dynamic control of the vehicle.

[0067] 4. Cloud-edge communication collaboration module, including:

[0068] Model incremental update: The cloud transmits the parameter increment of the policy model to the vehicle side to reduce communication overhead; the vehicle side fuses the increment on the basis of the existing model and updates the policy in a timely manner.

[0069] Selective feature upload: The vehicle side adaptively selects key environmental features to upload to the cloud according to feature importance and communication bandwidth, ensuring information quality while reducing communication load.

[0070] Federated learning aggregation: The cloud collects environmental features and model updates uploaded by multiple vehicles, aggregates and optimizes them through federated learning algorithms to improve the generalization performance of the policy model.

[0071] 5. System Management and Monitoring Module:

[0072] Responsible for coordinating the work processes of modules such as data collection, cloud training, vehicle-side execution, and communication transmission, ensuring the real-time performance and stability of the system.

[0073] Monitor the operating status of each part of the system, collect and analyze performance metrics such as decision-making latency, communication throughput, resource utilization, etc., to achieve fault diagnosis and performance optimization.

[0074] Provide a human-computer interaction interface, allowing users to configure system parameters, query system status, manually take over control, etc., to improve the usability and security of the system.

[0075] The above five modules have clear division of labor and cooperate with each other to form a complete cloud-edge collaborative autonomous driving control system. The data collection and preprocessing module provides a high-quality data foundation for the system; the cloud-based policy learning module uses powerful computing capabilities to train robust and general driving policies; the vehicle-side policy execution module efficiently executes policies under limited resources to control the vehicle to drive smoothly; the cloud-edge communication collaboration module realizes distributed collaborative optimization, balancing local personalization and global generalization; the system management and monitoring module ensures the efficient cooperation of each module, and the entire system operates safely, reliably, and with high performance.

[0076] The beneficial technical effects of the present invention are as follows: The cloud-edge collaborative autonomous driving control method based on the attention mechanism and GAN proposed by the present invention has achieved remarkable results in improving the performance of the autonomous driving system. Through end-to-end learning and the attention mechanism, this method can adaptively focus on key information in the driving environment, generate more accurate and smooth driving decisions, and effectively improve the scene adaptability and generalization ability of autonomous driving. At the same time, Generatove-Adversarial optimization further enhances the diversity and robustness of driving policies, enabling the autonomous driving system to cope with complex and changing actual traffic scenarios. In addition, the collaborative learning and heterogeneous computing architecture between the cloud and the vehicle side not only makes full use of the powerful computing capabilities of the cloud to extract general laws from massive data, but also gives play to the real-time processing capabilities of the vehicle side to execute personalized decisions. Supplementary mechanisms such as federated incremental learning and selective communication greatly compress the communication overhead and improve the real-time performance and scalability of the system. Combining the above technical innovations, the method of the present invention can significantly improve the safety, reliability, and generalization performance of the autonomous driving system, and accelerate the maturity and practical application of autonomous driving technology. Description of the Drawings

[0077] Figure 1 A structural diagram of an intelligent driving collaborative control system based on cloud-edge collaboration provided by the present invention. DETAILED DESCRIPTION

[0078] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0079] Using the data acquisition and preprocessing module, a large amount of autonomous driving data is collected in a real road environment, including multiple camera video images of the front, rear, left and right sides of the vehicle, as well as time series data of the vehicle's driving status such as speed, acceleration, steering angle, etc. In order to ensure the diversity of data, different scenes such as highways, urban roads, and rural roads are selected, as well as different conditions such as daytime, night, rainy and snowy weather. All data are aligned in time and space, that is, the sampling time of different sensors and the vehicle position are aligned. In order to ensure data quality, a series of preprocessing operations are performed on the original image, including:

[0080] (1) Image enhancement: adjust the brightness, contrast, saturation, etc. of the image to make it clearer and sharper.

[0081] (2) Noise removal: Use a median filter to remove the salt and pepper noise of the image, and use a Gauss filter to remove Gaussian noise.

[0082] (3) Distortion correction: Correct imaging distortion introduced by fisheye lenses, etc., and restore the linear perspective relationship of the image.

[0083] (4) Size normalization: Images of different resolutions are uniformly scaled to 256×256 to facilitate subsequent convolutional neural network processing.

[0084] The preprocessed image is denoted as Where h, w, and c are the height, width, and number of channels of the image, respectively.

[0085] Build a cloud-based Attention-GAN strategy model. In the model, the role of the Generator is to t and historical driving behaviora t-1 Generate safe and smooth next driving behavior It can be expressed as:

[0086]

[0087] Among them, θ g is the parameter of Generator; t Image from the vehicle camera I t , vehicle speed t , acceleration acc t etc., indicating the current environment; a t-1Control instructions for the accelerator pedal, brake pedal, steering wheel, etc. in the previous step. Since the original image I t contains a large amount of pixel redundancy, it is compressed into a compact feature map through a visual feature extractor f cnn (·):

[0088]

[0089] where h′, w′, c′ are the height, width, and number of channels of the feature map respectively, and h, w, c are less than or equal to h′, w′, c′ respectively. f cnn (·) represents the convolutional neural network feature extractor function, and a convolutional neural network with ResNet-18 as the backbone network is adopted.

[0090] To improve the adaptability of the policy model to complex traffic scenarios, an Attention mechanism is introduced in the Generator. Specifically, the feature map F t is first transformed into three parts: query / key / value through a 1×1 convolution:

[0091]

[0092] where is a learnable transformation matrix. Then, the attention weights of the query and the key are calculated:

[0093]

[0094] where c″ is a scaling factor used to control the magnitude of the dot product. Finally, the value feature is weighted with the attention weight A t to obtain the attention-enhanced visual feature:

[0095]

[0096] This process enables the model to focus on the similarity of the key at other positions according to the current query and aggregate the relevant value features. Compared with the original feature F t , the attention feature highlights the targets crucial for driving decisions, such as vehicles, pedestrians, traffic signs, etc.

[0097] Next, the attention feature is expanded into a vector (representing the new set after image dimension scaling here), and concatenated with the state features such as vehicle speed v t , acceleration acc t etc., and then fed into an LSTM cell to model the temporal dependence:

[0098]

[0099] where h t and c t are the hidden state and memory state of the LSTM respectively; W o and b o are the parameters of the output layer. The LSTM can learn the correlation of driving behavior in the time dimension, making the generated trajectory smoother and more coherent. Finally, the output h t of the LSTM is mapped to a normalized driving behavior through a tanh activation function

[0100] The goal of constructing the Discriminator is to determine whether a driving behavior sequence is from a real driver or the Generator. Its input is a sequence of traffic environment states and the corresponding behavior sequence and the output is the authenticity score of the behavior sequence

[0101]

[0102] where D(·) represents the Discriminator function and θ d is the parameter of the Discriminator. Specifically, the same visual feature extractor f cnn (·) as the Generator and the attention mechanism can be used to process the image sequence to obtain a sequence of attention feature maps Then it is concatenated with the behavior sequence and fed into a behavior recognition network composed of multiple one-dimensional convolutional layers and LSTM layers:

[0103]

[0104] where Conv1D(·) represents the convolution operation in the time dimension; σ(·) is the Sigmoid function, which compresses the output to the range of (0, 1). The idea of the Discriminator is that the behavior of real drivers usually adapts to the traffic environment, such as decelerating when encountering a vehicle or stopping at a red light, while the behavior generated by the Generator at the initial stage may not fit the environment well. Therefore, the behavior recognition network can learn an environment-behavior matching pattern to distinguish real and generated behaviors.

[0105] The Generator and the Discriminator are optimized through adversarial training, and its objective function is:

[0106]

[0107] The first term is the log probability of the real behavior by the Discriminator, and the second term is the log probability of the generated behavior by the Discriminator. The objective is as follows: The Generator generates more realistic behaviors to deceive the Discriminator, while the Discriminator distinguishes real and generated behaviors as accurately as possible. The two alternate during the training process and balance each other, ultimately enabling the Generator to learn to imitate real drivers.

[0108] The specific training process is as follows:

[0109] Input: historical environmental state and behavior data s, a,

[0110] Initialize the parameters θ of the Generator and the Discriminator g , θ d ;

[0111] 1. Sample a batch of real data s, a m ~p data , and random noise z m ~p z ;

[0112] 2. Generate a batch of fake behaviors

[0113] 3. Calculate the losses of the Discriminator on real and fake behaviors:

[0114]

[0115] 4. Update the parameters of the Discriminator:

[0116] 5. Sample a batch of random noise z m ~p z ;

[0117] 6. Generate a batch of fake behaviors

[0118] 7. Calculate the loss of the Generator:

[0119] 8. Update the parameters of the Generator:

[0120] Among them, α is the learning rate. The goal of the Discriminator is to maximize the sum of the logarithmic probability of real behaviors and the negative logarithmic probability of false behaviors, while the goal of the Generator is to minimize the negative logarithmic probability of false behaviors. They are opponents of each other, constantly gaming in iterations, and finally reaching a dynamic balance.

[0121] Due to the limited computing and storage resources of in-vehicle devices, a lightweight execution network is designed on the vehicle side to run the driving strategies learned in the cloud and timely feedback the vehicle state to the cloud.

[0122] The core of the execution network is a policy network π, which takes the current environmental state s t and the previous behavior a t-1 as inputs to generate the optimal driving behavior distribution μ t :

[0123] μ t = π(s t , a t-1 ; θ π );

[0124] Among them, θ π is the parameter of the policy network. μ t is a multi-dimensional Gaussian distribution, whose mean represents the optimal behavior and the variance represents the uncertainty of the behavior. Specific driving instructions a t can be sampled from this distribution, such as:

[0125]

[0126] The policy network adopts the same structure as the Generator in the cloud, including visual feature extraction, attention mechanism, and LSTM sequence modeling components. However, considering the resource constraints on the vehicle side, a series of optimizations are carried out on it:

[0127] 1) Compress the visual feature extractor and adopt lightweight backbone networks such as MobileNet and ShuffleNet.

[0128] 2) Reduce the resolution of the attention feature map and sacrifice spatial details to reduce the computational amount.

[0129] 3) Replace LSTM with GRU to process temporal dependencies, with fewer parameters and faster computation.

[0130] 4) Perform channel pruning and binary quantization on the policy network to further compress the model size.

[0131] The optimized policy network can run smoothly on in-vehicle devices, achieving a decision-making delay of milliseconds. Meanwhile, a lightweight environment perception module is deployed on the vehicle side to extract key information of the driving environment and upload it to the cloud, such as:

[0132] The positions and speeds of pedestrians and vehicles;

[0133] The categories and states of road signs and traffic lights;

[0134] The geometric shapes of lane lines and stop lines.

[0135] These information are compressed into a compact feature vector x t , and are uploaded to the cloud in real time through a wireless network for continuously optimizing the cloud policy model. Through the training of the cloud policy in a massive environment and the adaptation of the vehicle-side policy in a personalized environment, the complementary improvement of the capabilities of both can be achieved.

[0136] To reduce the communication load, instead of directly synchronizing all the model parameters trained in the cloud to the vehicle side in full, an incremental update method is adopted. Let the cloud model parameters after the k-th update be θ k , then the increment Δθ k is:

[0137] Δθ k = θ k - θ k-1 ;

[0138] Only Δθ k is sent to the vehicle side, and the vehicle side obtains the latest model by accumulating the first k increments:

[0139]

[0140] Generally, the model tends to converge in the later stage of training, and the parameter update amplitude will gradually become smaller. Therefore, incremental transmission can significantly reduce the communication volume. In addition, sparsification and quantization are also performed on the increment to further compress its size:

[0141] Sparsification: Only transmit the increment components with amplitudes greater than the threshold λ, and the rest of the components are approximated as 0. That is

[0142] Quantization: Each component of the increment is encoded with an n-bit integer, sacrificing a small amount of precision in exchange for a smaller bandwidth.

[0143] After sparse quantization, the size of the increment can be compressed to one-tenth or even less of the original, greatly improving the communication efficiency.

[0144] Selective upload of environmental features. The environmental feature x t perceived by the vehicle side is a high-dimensional vector, and uploading all the content will occupy a large amount of bandwidth. But xt Often, only a small part of the dimensions are highly relevant to driving decisions, and the value of most dimensions is low. Therefore, the vehicle end maintains a feature importance measure ρ ∈ [0, 1] based on historical data statistics d :

[0145]

[0146] where d is the feature dimension, x i is the i-th dimensional feature, a is the driving behavior, and MI(·, ·) is the mutual information, which measures the correlation between two variables. ρ statistically calculates the correlation degree between each feature dimension and the behavior from historical data, and the score of the dimension with the highest importance is 1

[0147] When uploading, the vehicle end adaptively selects some dimensions to upload according to the importance measure ρ and the current communication bandwidth B. Let the upload ratio be γ ∈ (0, 1], then the uploaded feature subset is:

[0148]

[0149] That is, select the dimensions ranked top in importance. The size of γ is positively correlated with the bandwidth B, and the larger the bandwidth, the higher the upload ratio. The uploaded feature z t is:

[0150] z t = m t ⊙ x t ;

[0151]

[0152] where ⊙ is the element-wise multiplication, and m t is the mask vector. The unselected dimensions are directly set to zero and do not occupy the communication bandwidth. Through this selective upload mechanism, the vehicle end can transmit the most valuable environmental information to the cloud with the least communication volume

[0153] To balance the computational loads of the cloud and the vehicle ends and achieve distributed cooperation in model training and inference, the present invention also adopts the framework of federated learning. Suppose there are N vehicle i-end users participating in collaborative training, the cloud model is θ, and the local model of the i-th vehicle end is θ i , and the local dataset is The goal of federated learning is:

[0154]

[0155] where f(θ) is the global objective function, and F i (θ) is the objective function of the i-th vehicle end on the local data, that is, the empirical risk:

[0156]

[0157] l(θ; s, a) is the loss function of the model θ taking action a in the environment s. It can be seen that the global objective is the weighted average of all vehicle-end objectives, and the weights are proportional to the number of local samples. Intuitively, vehicle-ends with more data have a greater proportion in the global model.

[0158] Federated learning optimizes the global model through the following steps:

[0159] 1) The cloud broadcasts the latest model θ t to all online vehicle-ends;

[0160] 2) Each vehicle-end fine-tunes the model on its local data for several rounds to obtain the updated model θ i,t+1 ;

[0161] 3) The vehicle-end uploads the model increment Δθ i,t+1 = θ i,t+1 - θ t to the cloud;

[0162] 4) The cloud weights and averages all the increments according to the number of samples to obtain the global increment:

[0163]

[0164] 5) The cloud updates the global model: θ t+1 = θ t + Δθ t+1 .

[0165] The above process is iterated so that the cloud model gradually absorbs the knowledge of all vehicle-ends and obtains wide applicability. On the other hand, each vehicle-end model continues to perform personalized learning based on the cloud model to adapt to its respective driving style. The frequent interactive learning between the cloud and vehicle-ends synchronously improves the capabilities of both.

[0166] Federated learning can complete large-scale collaborative learning with relatively low communication costs while protecting user privacy. However, since the data distributions of various vehicle-ends may vary greatly, heterogeneity will affect the convergence speed and effect of the global model. Therefore, based on the original federated averaging algorithm, the present invention introduces the following improvements:

[0167] Weighted aggregation: Different aggregation weights are assigned to different vehicle-ends according to the quality of the vehicle-end data and the credibility of the model. Vehicle-ends with high quality and high credibility have large weights.

[0168] Dynamic adjustment of the learning rate: According to the convergence situation of the model, the learning rates of the cloud and vehicle-ends are adaptively adjusted. A large learning rate is used for rapid convergence in the initial stage of training, and a small learning rate is used for fine-tuning in the later stage.

[0169] Control the update frequency: Reasonably set the update frequencies of the cloud and vehicle sides to achieve a balance between communication efficiency and learning effect. If the frequency is too high, the communication overhead is large; if the frequency is too low, the convergence speed will be affected.

[0170] Asynchronous update: Allow some vehicle sides to delay updates or go offline. As long as a certain coverage rate is reached, it can trigger cloud aggregation, avoiding the synchronization overhead of waiting for all vehicle sides.

[0171] The improved federated learning framework achieved better results than traditional centralized learning and single-vehicle-side learning in simulation experiments and demonstrated good robustness.

[0172] According to the above plan, in the embodiment, a large-scale autonomous driving dataset was collected in a real road environment, including various scenarios such as highways, urban areas, rural areas, and mountainous areas, and various weather conditions such as day, night, rain, snow, etc., covering typical driving environments in China.

[0173] Randomly divide the data into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set is used to train the model, the validation set is used for model selection and hyperparameter search, and the test set is used to evaluate the generalization performance of the model.

[0174] To objectively measure the model effect, the following evaluation metrics are adopted:

[0175] 1) Mean Absolute Trajectory Error (MATE):

[0176] where N is the number of test samples, T is the trajectory length, and y i,t are the predicted value and the true value of the i-th sample at time t, respectively. MATE measures the average difference between the predicted trajectory and the true trajectory, and the smaller it is, the more accurate the prediction.

[0177] 2) Weighted Trajectory Error (WTE):

[0178] where w t is the weight at time t. WTE introduces a time weight on the basis of MATE, highlighting the importance of short-term prediction accuracy because drivers are more concerned about the immediate driving trajectory. The weights for the first 1 second, 1 - 3 seconds, and after 3 seconds are set to 0.6, 0.3, and 0.1, respectively.

[0179] 3) Success rate:

[0180] where 1(·) is the indicator function, which takes 1 when the i-th test sample successfully completes trajectory tracking and has no collision, and 0 otherwise. The success rate intuitively reflects the proportion of safe trajectories planned by the model.

[0181] 4) Average Decision Time (ADT):

[0182] where Δt i,t is the decision-making time consumption of the i-th sample at time t. ADT examines whether the model can meet the real-time requirements from the perspective of computational efficiency.

[0183] The above indicators reflect the model performance from multiple aspects such as trajectory similarity, safety, and real-time performance, and can comprehensively evaluate the advantages and disadvantages of different models. In addition, indicators reflecting driving smoothness such as the number of collisions, the number of hard brakes, and the number of sharp turns are also counted.

[0184] The Attention-GAN model was implemented in Python under the PyTorch deep learning framework and trained on a workstation equipped with 8 Tesla V100 GPUs. The hyperparameters for training are as follows:

[0185] Both the Generator and the Discriminator use the Adam optimizer, with learning rates of 0.0001 and 0.0004 respectively, and a weight decay coefficient of 0.0001.

[0186] The feature dimension of Attention is 128.

[0187] In the Discriminator, the convolutional kernel sizes are 32, 64, 128, 256 in sequence, and the number of channels is 128, 256, 512, 1024 in sequence.

[0188] The Batch size is 64, and training is carried out for 100 epochs until the performance of the validation set saturates.

[0189] At the end of each epoch, the MATE and WTE metrics are evaluated on the validation set, and the model parameters with the best performance are saved for testing. To prevent overfitting, strategies such as L2 regularization, Dropout, and early stopping are adopted.

[0190] Next, the trained cloud model is deployed to the vehicle side, and cloud-edge synchronization is performed through an incremental update mechanism. The vehicle-side execution network is deployed on the NVIDIA DRIVE PX2 autonomous driving platform with the Linux operating system. The aforementioned various sensors are installed on each test vehicle, and the vehicle travels on real roads in multiple cities in China to evaluate the model performance online.

[0191] To demonstrate the superiority of this model, the following comparative experiments were set up:

[0192] 1) Cloud-only: Only train the model in the cloud, and all vehicles share the same model without personalized learning.

[0193] 2) Edge - only: Each vehicle trains its own model independently without interacting and learning with other vehicles.

[0194] 3) FedAvg: The traditional federated averaging algorithm, where the cloud and vehicle sides use the same global model and the weights are aggregated equally.

[0195] 4) FedAttGAN: The Attention - GAN federated learning architecture proposed in this invention.

[0196] Table 1: Quantitative comparison of different models on the test set

[0197] Model MATE(m) WTE(m) Success Rate (%) ADT(ms) Cloud-only 1.35 1.42 88.5 25 Edge-only 1.19 1.28 91.7 76 FedAvg 0.95 0.97 94.2 43 FedAttCAN 0.76 0.81 96.8 51

[0198] Table 1 shows the quantitative metrics of each model on the test set. FedAttGAN is significantly better than the other three schemes. Specifically:

[0199] 1) Cloud - only is difficult to adapt to all vehicles due to ignoring the individual differences between vehicles, and has the largest mean average trajectory error (MATE). However, due to the largest model scale and the most abundant computing resources in the cloud, the average decision time (ADT) is the shortest.

[0200] 2) Edge - only can adapt to its own driving style and habits through independent learning of each vehicle, so the trajectory error is smaller than that of Cloud - only. However, limited by the computing power of in - vehicle devices, its decision time is longer than that of Cloud - only, and the success rate is slightly lower.

[0201] 3) FedAvg achieves a better balance between Cloud - only and Edge - only, absorbs the generalization of the cloud and the personalization of the vehicle side, and all indicators are improved. This shows that federated learning can effectively integrate the experiences of different vehicles and obtain a more robust driving strategy.

[0202] 4) Based on FedAvg, FedAttGAN further improves the end - to - end modeling ability of perception, decision - making, and control through Attention and GAN, and improves the federated aggregation algorithm, making the trajectory error and success rate reach new highs while maintaining a relatively low decision time.

[0203] In summary, in view of the three major problems of perception, decision-making, and control faced in the field of autonomous driving, the present invention proposes a cloud-edge collaborative control method based on Attention-GAN. It innovatively combines end-to-end learning and federated learning, designs the collaborative optimization of the driving strategy between the cloud GAN and the vehicle-side execution network, and introduces an attention mechanism to model the interaction between the vehicle and the environment. Large-scale experiments on real road data show that the proposed method can significantly improve the trajectory tracking accuracy, scene adaptability, and driving safety of autonomous driving, laying a technical foundation for the mature application of driverless driving.

[0204] The present invention also provides a computing device, which may include: a processor, a communications interface, a memory, a display screen, and an input device. Among them, the processor, the communications interface, and the memory complete their mutual communication through a communication bus.

[0205] The processor is used to provide computing and control capabilities. The processor can call the logical instructions in the memory. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. When the computer program is executed by the processor, it implements the methods in the above-mentioned various embodiments; the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communications interface is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computing device, or an external keyboard, touchpad, or mouse, etc.

[0206] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0207] In an embodiment of the present invention, a non-transitory computer-readable storage medium is further provided. The non-transitory computer-readable storage medium stores server instructions, and the computer instructions cause the computer to execute the methods provided in the above embodiments. For a computer-readable storage medium provided in the above embodiments, its implementation principle and technical effects are similar to those of the above method embodiments, and will not be elaborated here.

[0208] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0209] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0210] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

Claims

1. An intelligent driving collaborative control method based on cloud-edge collaboration, characterized in that The method includes the following steps: Step 1: Collect and preprocess intelligent driving-related data in the cloud-edge system; Step 2: Build an Attention-GAN policy model in the cloud; Step 3: Conduct Generative-Adversarial training in the cloud; Step 4: Deploy the execution network on the vehicle side; Step 5: Design the cloud-edge communication mechanism.

2. The intelligent driving collaborative control method based on cloud-edge collaboration according to claim 1, wherein, The said Step 1 includes: Step 1.1: Collect autonomous driving data in the real road environment, including sequential information such as in-vehicle camera images, vehicle speed, acceleration, etc.; Step 1.2: Perform preprocessing on the original images for enhancement, denoising, and correction, and unify the size to h×w×c, where h represents the image height, w represents the image width, and c represents the number of image channels; Step 1.3: Divide the data set into a training set, a validation set, and a test set in the ratio of 8:1:

1.

3. The intelligent driving collaborative control method based on cloud-edge collaboration according to claim 1, wherein, The said Step 2 specifically includes: Step 2.1: Construct Generator G, which generates false driving behaviors based on the current environmental state s t and historical actions a t-1 ​ where θ g is a parameter of the Generator; Step 2.2: Introduce a visual feature extractor to extract the compact features of the images: Among them, F t represents the extracted image features, and f cnn (·) represents the convolutional neural network feature extractor function, I t is the input original image, h′ represents the height of the feature map, w′ represents the width of the feature map, and c′ represents the number of channels of the feature map; Step 2.3: Design an attention mechanism to calculate the weight A through query, key, and value t , and obtain the attention feature by weighting the feature map Among them, A t represents the attention weight matrix, represents the weighted attention feature, softmax(·) represents the softmax normalization function, Q t represents the query matrix, K t represents the key matrix, V t represents the value matrix, and c″ is the scaling factor; Step 2.4: Input the attention feature and the status information into the LSTM to model the temporal dependence and generate the final driving behavior Step 2.5: Build a Discriminator to distinguish real and generated behavior sequences: Among them, represents the authenticity score of the behavior sequence, and D(·) represents the Discriminator function. represents the time series data of the environmental state and behavior, T represents the time series length, and θ d is the parameter of the Discriminator.

4. The intelligent driving collaborative control method based on cloud-edge collaboration according to claim 1, wherein, The said Step 3 specifically includes: Step 3.1: Train the Discriminator to maximize the log probability of real behaviors and the negative log probability of false behaviors: Among them, represents the loss function of the Discriminator, m is the batch number, s i represents the i-th true environmental state, a i represents the i-th true behavior, represents the i-th false behavior generated by the Generator; Step 3.2: Train the Generator to minimize the negative log probability of false behaviors: Among them, represents the loss function of the Generator; Step 3.3: Alternately optimize the Generator and the Discriminator until the model converges.

5. The intelligent driving collaborative control method based on cloud-edge collaboration according to claim 1, characterized in that The said Step 4 specifically includes: Step 4.1: Design a lightweight policy network π, input the environmental state s t , the historical action a of the previous step t-1 , and output the optimal driving behavior distribution μ t : μ t = π(s t , a t-1 ; θ π ); Among them, μ t represents the generated behavior distribution, π(·) represents the policy network function, and θ π is the parameter of the policy network; sample the driving instruction a t from μ t ; Step 4.2: Use model compression technology to optimize the policy network and control the computational overhead; Step 4.3: Develop an in-vehicle environment perception module to extract key driving information and upload it to the cloud.

6. The intelligent driving collaborative control method based on cloud-edge collaboration according to claim 1, characterized in that, The said Step 5 specifically includes: Step 5.1: Synchronize the cloud and vehicle models in an incremental update manner, that is, only transmit the increment Δθ k To the vehicle side: Δθ k = θ k - θ k-1 ; Among them, Δθ k represents the increment of the model parameters for the k-th update, θ k represents the cloud model parameters after the k-th update, θ k-1 represents the cloud model parameters after the (k - 1)-th update, represents the vehicle-end model parameters after the k-th update; and sparsely quantizes and compresses the increment; Step 5.2: The vehicle side selectively uploads environmental features according to the communication bandwidth B and the feature importance metric ρ: Among them, is the subset of feature dimensions uploaded, and ρ i represents the importance measure of the i-th dimensional feature, γ ∈ (0, 1] is the upload ratio, and d represents the total dimension of the features; Step 5.3: Design a federated learning framework, and the cloud aggregates the local updates of different vehicles to optimize the global policy: Among them, f(θ) is the global objective function, F i (θ) is the local objective function of the i-th vehicle, N is the total number of vehicles participating in federated learning, represents the local dataset of the i-th vehicle, represents the datasets of all vehicles, is the loss function for the model θ to take action a under the environment s.

7. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the methods described in claims 1-6.

8. A computing device, characterized in that: It includes one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described in claims 1-6.

9. A cloud-edge collaborative intelligent driving collaborative control system for implementing the method according to any one of claims 1-6, characterized in that, The system includes the following modules: a cloud-edge collaborative intelligent driving data collection and preprocessing module, a cloud policy learning module, a vehicle-side policy execution module, a cloud-edge communication collaboration module, and a system management and monitoring module; The cloud-edge collaborative intelligent driving data collection and preprocessing module collects multi-source heterogeneous data required for autonomous driving, and conducts data cleaning, annotation, enhancement, and preprocessing, and divides the data set into a training set, a validation set, and a test set for model learning and evaluation; The cloud policy learning module includes a Generator based on the attention mechanism and a Discriminator based on a convolutional neural network; The vehicle-side policy execution module includes: The lightweight policy network compresses and deploys the policy model trained in the cloud to in-vehicle devices, receives the environmental status in real time, and outputs driving behavior decisions; The environmental perception unit uses in-vehicle sensors to collect traffic environment information in real time, extracts and uploads key features to the cloud; The policy execution unit executes the driving decisions generated by the policy network and controls the vehicle actuators; The cloud-edge communication and collaboration module includes: model incremental update, selective feature upload, and federated learning aggregation; The system management and monitoring module is responsible for coordinating the work processes of data collection, cloud training, vehicle-side execution, and communication transmission, and provides a human-machine interaction interface.

Citation Information

Cited By

  • Intelligent driving cooperative control terminal based on visual language action model

    CN121133732A

  • Intelligent driving cooperative control terminal based on visual language action model

    CN121133732B

  • Intelligent networked automobile multi-mode fusion automatic driving decision-making platform

    CN122090418A