An intelligent vehicle scenario understanding method based on enhanced edge data learning
By extracting and processing edge scene data, combining YOLO and graph neural network technology, a scenario understanding data process is built, and the big model is fine-tuned and trained, the existing big model is solved, and the existing big model performs poorly in understanding and prediction of complex scenes is realized, and the efficient understanding and prediction capabilities of smart cars in complex scenarios are achieved.
Patent Information
- Application Number
- CN202510192921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing large models are difficult to effectively utilize complex, edge driving scenario data in the understanding of smart car scenarios, resulting in poor performance in understanding complex scenarios and predicting future changes.
By analyzing the road collection data, edge scene data is extracted; structured templates are built based on the scene ontology, combined with YOLO architecture and graph neural network, test scene pictures are deconstructed, element characteristics in the scene are extracted and detailed descriptions of scenes are improved; based on future scene information extraction and main car response measures annotation, a data process from scene cognition to functional applications is built; open source large model architecture is used for model fine-tuning training, and the model's understanding of complex and rare driving conditions is enhanced.
The understanding of smart car scenarios enhanced based on edge data learning is realized, which improves the model's understanding and interpretability in complex and edge scenarios, and ensures the accuracy of future state prediction and generalization of new and old tasks.
Smart Images

Figure CN119693906B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving testing, and specifically relates to an intelligent vehicle scene understanding method based on enhanced edge data learning. Background Art
[0002] With the wide application of deep learning, machine learning, and complex algorithms in intelligent vehicle systems, vehicles can demonstrate capabilities beyond human drivers in aspects such as autonomous driving, road decision-making, and risk warning. However, behind this high level of intelligence lies the "black box" effect: even though intelligent vehicle systems can effectively execute various driving tasks, their decision-making processes often lack transparency and comprehensibility, which is crucial for the safety and reliability of intelligent vehicle algorithms. Therefore, developing intelligent vehicle algorithms that can self-explain, are easy to understand and debug has become a major issue faced by global scientific researchers and technology developers.
[0003] Scene understanding actively perceives, analyzes, and interprets driving scenes through a computer. For intelligent vehicle algorithms, it can clearly display the external perception environment and internal decision-making logic to users, thereby improving the interpretability of the algorithms and enhancing users' trust and acceptance of intelligent vehicles. At the same time, through scene understanding, intelligent vehicles can prompt users when the scene is abnormal so that users can take timely measures.
[0004] Large models use massive amounts of data for artificial intelligence model training and have powerful capabilities in image and natural language understanding, showing significant potential advantages in intelligent vehicle scene understanding. However, due to the long-tail effect of natural driving data, the datasets used for large model training often only contain basic information. Therefore, existing large models can only understand basic scene features and are difficult to be directly used for complex and edge intelligent vehicle scene understanding tasks. Edge scene data has rich and diverse scene features. How to efficiently apply it to the large model training process and expand the application fields of large models is a scientific problem that urgently needs to be solved. Specifically, it is manifested as the following three challenges: (1) How to obtain and utilize those rare, complex, and low-proportion edge driving scene data in the conventional training set; (2) How to design effective data augmentation strategies so that large models can learn rich scene features from limited edge scene data; (3) How to optimize for complex edge scene understanding tasks of intelligent vehicles while ensuring the original performance of large models, thereby promoting the practical application process of large models in the field of intelligent vehicles. Summary of the Invention
[0005] To solve the above problems, the present invention provides an intelligent vehicle scene understanding method based on enhanced edge data learning. By analyzing the road acquisition data, edge scene data is extracted; a structured template is constructed based on the scene ontology, and combined with the YOLO architecture and graph neural network, it can effectively deconstruct the test scene pictures, extract the element features in the scene, and improve the description of scene details; based on the extraction of future scene information and the annotation of the host vehicle's response measures, it can effectively build a complete data flow from scene cognition to functional application, and construct a model training data set; based on the open-source large model architecture, model fine-tuning training is carried out, and finally, the intelligent vehicle scene understanding based on enhanced edge data learning is realized, which can be used for understanding and cognition in complex and edge scenes, and improve the interpretability of intelligent vehicle algorithms.
[0006] The technical solution of the present invention is described in conjunction with the accompanying drawings as follows:
[0007] An intelligent vehicle scene understanding method based on enhanced edge data learning, comprising the following steps:
[0008] Step 1: Based on the real vehicle vision acquisition platform, collect the driving conditions of natural driving in the urban area;
[0009] Step 2: Construct a structured template based on the scene ontology; use the YOLO vision processing algorithm for object detection and semantic label classification; utilize the graph neural network to fill and map the scene image into a predefined scene template to complete the conversion from image to structured data;
[0010] Step 3: Develop and apply specific metrics to automatically screen out challenging edge scenes. For cases where the classification accuracy is lower than the set threshold, introduce a secondary review to improve the data quality;
[0011] Step 4: Match the possible future changes according to the scene semantic labels, and mark the actions that the vehicle should take under the current situation; perform preprocessing on the original image and encode it into a vector format for constructing a training data set, with the scene vector data as the input and the structure template representation, future change data, and host vehicle response measures as the output;
[0012] Step 5: Use the mature qwen-vl architecture and fine-tune it for the edge scene data set, so that the model can learn to extract features from the image and make accurate future predictions; in order to maintain the generalization of the trained model, add the cross-entropy of the outputs of the two models to the loss function to quantify the difference between the outputs of the original model and the new model, and ensure that the learning of the new model does not deviate significantly from the performance of the original model; regularly verify the model performance and adjust the training parameters to ensure the generalization ability of the model until a satisfactory performance level is reached;
[0013] Step 6: Select an open-source large model as a benchmark, compare its performance in basic cognition, future change prediction, and vehicle behavior analysis, and evaluate the practical application value of the model; analyze the performance differences between the new model and the original model on the same test set to confirm whether the new model has successfully retained the original task understanding ability.
[0014] Further, the specific method of the above Step 1 is as follows:
[0015] S11: Build a test vehicle data acquisition platform based on the Horizon Vision Suite;
[0016] S12: Capture diverse driving scenarios in real time: Through the camera devices installed on the intelligent vehicle, record the visual information of the vehicle during urban driving in real time;
[0017] S13: When selecting the vehicle driving area, ensure that the data covers a variety of driving scenarios.
[0018] Further, the specific method of the above Step 2 is as follows:
[0019] S21: Construct a scene structured template. Based on ontology, abstract and summarize the driving scenarios into a series of structured templates, including a term layer, an attribute layer, and an association layer. By annotating according to this template, the spatial relationships and interaction rules between different elements can be defined;
[0020] S22: Classify the scene semantic labels. Use the visual processing algorithm YOLO to perform in-depth analysis on the images obtained in the natural driving data collection stage;
[0021] S23: Fill and transform the images. Map the scene elements with semantic labels to the preset scene structured template. Based on the graph neural network, transform the originally unstructured scene images into a structured representation with clear semantic meanings and positional relationships, and transform the natural driving data into a form that can be efficiently utilized by the large model;
[0022] Further, the specific method of the above S22 is as follows: Use the YOLO-v8 architecture to process the collected data, extract the terms and basic attributes in the ontology architecture diagram, locate and identify each object in the image, assign corresponding semantic labels to each object, and output the class probability, center point coordinates (x, y), width (w), and height (h) of each type of bounding box, and convert them into bounding box representations.
[0023] Further, the specific method of the above S23 is as follows:
[0024] S231: Take each detected object as a node and add edges to express the spatial relationships or semantic relationships between the objects;
[0025] S232. Create a graph neural network model to capture the interactions between nodes and propagate context information; define a node embedding layer to encode the original features of each object, i.e., the bounding information output by YOLO, into a high-dimensional vector representation; design a graph convolutional layer and a message passing mechanism suitable for structured representation learning, enabling the network to learn the complex relationships between nodes.
[0026] S233. Train the graph neural network model to learn and infer deeper semantic relationships and structured information from the graph structure; after training, input the YOLO detection results into the graph neural network model, and the model will output updated node representations that can be mapped to the concepts and their relationships in the ontology framework, thereby forming a structured representation.
[0027] Furthermore, the specific method of step three is as follows:
[0028] S31. Edge scenario definition: There is no clear definition of edge scenarios in the academic community. Considering that edge scenarios pose challenges to existing intelligent algorithms, the present invention defines edge scenarios as those situations that pose challenges to current intelligent algorithms, specifically including the following two parts:
[0029] ① Uncertain identification: The target detection confidence output by the YOLO algorithm is lower than the set threshold (set to 0.7), indicating that the model is uncertain about whether the detected object belongs to a certain category.
[0030] ② Poor structuring: After processing by the graph neural network (GNN), the node classification accuracy fails to reach the predetermined standard (set to 90%), meaning that the elements in the image are not correctly classified and associated.
[0031] S32. Automatic extraction of edge scenarios: Through the above-mentioned scene marginality judgment index system, automatically screen out those edge driving scenario images that cannot be accurately identified or have poor structuring effects. These edge scenarios may include rare road conditions, special traffic events, views under extreme weather conditions, etc.
[0032] S33. Automatic data improvement: For the edge scenario images automatically extracted, due to the poor performance of the target detection algorithm and the scene structuring algorithm, it is difficult to generate accurate and complete scene structured data. Therefore, the system introduces an automatic annotation technology to improve these data, ensuring that the model can learn and adapt to complex scenarios more effectively. Through weak supervision learning methods, perform secondary annotation classification on various types of targets in the image. Through the iterative learning process, the system can gradually improve the understanding accuracy of the elements in complex scenarios, ensure that all elements of each frame of the image are marked, and thus improve the knowledge base of the model when processing such complex scenarios.
[0033] Furthermore, the specific method of step four is as follows:
[0034] S41. Predict the matching scenario information; analyze the original collected video or image sequence based on the scenario semantic tags, extract the state changes of the scenario within a continuous time period (the single-sequence time length is 5s), and construct the scenario prediction information;
[0035] S42. Label the behavior of the host vehicle; for the current scenario, combine the road traffic rules and the principles of safe driving, convert the above rules into prompt words through a pre-trained large language model, so as to automatically generate the driving strategies or actions that the host vehicle should take in a specific scenario; in addition, correct the generated strategies based on the above rules and delete the behaviors that do not conform to the rules;
[0036] S43. Preprocess the image, perform a series of preprocessing operations on the collected original image data, specifically including four steps: color space conversion, image cropping, normalization, and size scaling;
[0037] S44. Feature vectorization, convert the preprocessed image into a numerical feature vector that can be understood by a computer, use a convolutional neural network to extract image features, and map the image features from a high-dimensional pixel space to a low-dimensional feature space;
[0038] S45. Construct a dataset, combine the above processing results to construct a complete dataset containing scenario information; each sample includes the image data preprocessed and encoded as a feature vector, the corresponding scenario semantic label, the predicted future state change data, and the label of the ideal response measures that the host vehicle should take in the current scenario.
[0039] Furthermore, the specific method of step five is as follows:
[0040] S51. Select and initialize the model architecture: select the qwen-vl open-source large model architecture; initialize a new model based on the pre-trained qwen-vl weights;
[0041] S52. Apply the dataset: use the dataset constructed in step four that includes the prediction of edge scenario information for training to enhance the model's understanding of complex and rare driving conditions;
[0042] Using the LoRA (Low-Rank Adaptation) method, only the part of the model used to process new tasks is adjusted, especially the low-rank decomposition matrix in the fully connected layer. This method allows the network to focus on learning features related to new tasks without changing the original pre-trained weights. According to the format of the image feature vector, the dimension of the input layer is adjusted to be consistent with the feature vector in the aforementioned preprocessing stage. To maintain the generalization of the model, a cross-entropy term is added to the standard loss function to quantify the difference between the outputs of the original model and the new model. This ensures that the learning of the new model does not deviate significantly from the performance of the original model while improving the accuracy of future state prediction;
[0043] S53. Performance Monitoring and Adjustment: During the entire training cycle, regularly (every two rounds of training) evaluate the performance of the model on an independent validation set, and pay attention to the performance metrics of the model on specific tasks (such as scene understanding accuracy, future state prediction error, and rationality score of the host vehicle's response measures). Dynamically adjust the training strategy according to the evaluation results. Use a cosine annealing learning rate scheduler to automatically adjust the learning rate according to the performance on the validation set. The learning rate gradually decreases from the initial value to the minimum value and then increases again, forming a periodic change. According to the overfitting situation, appropriately adjust the L2 regularization coefficient to prevent the model from overfitting the training data. The adjustment process gradually increases by 10% until the overfitting phenomenon disappears;
[0044] Among them, the L2 regularization coefficient is the decay of the loss weight, and the calculation method of the loss function is as follows:
[0045] In the formula, is the original loss function, is the L2 regularization coefficient, is the weight coefficient of the model.
[0046] Further, the specific method of step six is as follows:
[0047] S61. Apply the Test Model: Apply the qwen-vl model fine-tuned and trained in step five and other typical open-source large image recognition models, such as EfficientNet and ResNet, to the edge scenario dataset constructed and processed in steps one to four for scene understanding analysis;
[0048] S62. Cognitive Evaluation of Scenarios: Compare the performance of each model in basic scene recognition, and judge whether each model can accurately identify vehicles, pedestrians, traffic signs, road conditions in the scene, and whether it can grasp the overall layout and structural characteristics of the scene; The specific indicators are designed as: perception category accuracy;
[0049] S63. Predict and Compare Future Changes: Analyze the ability of each model to predict future changes in scenarios, and compare the accuracy of different models. The specific indicators are designed as: the prediction accuracy of moving objects and the assessment accuracy of collision risks.
[0050] S64. Simulate and Evaluate Vehicle Behavior: For each scenario, evaluate the effect of each model in simulating the reasonable behavior that an autonomous vehicle should take, and reflect the decision-making response in complex road conditions (whether it complies with traffic regulations and safe driving guidelines). The specific indicator is designed as: the rationality of expected behavior.
[0051] S65. Comprehensively Evaluate Model Performance: Summarize the above evaluation results, compare the advantages and disadvantages of different models in edge scenario understanding and prediction tasks from multiple dimensions, find the most suitable large model for intelligent vehicle application scenarios, and at the same time identify the deficiencies of the models to provide directions for further optimizing model performance. Set the weights of each indicator to 0.4, 0.3, and 0.3, and perform weighted averaging on the scores of each indicator.
[0052] The beneficial effects of the present invention are as follows:
[0053] 1. Based on the scenario ontology, the present invention constructs a structured template, combines the YOLO architecture and the graph neural network, and can effectively deconstruct the test scenario pictures, extract the element features in the scenarios, and improve the detailed description of the scenarios.
[0054] 2. Based on the extraction of future scenario information and the annotation of the host vehicle's response measures, the present invention can effectively build a complete data flow from scenario recognition to functional application and construct a model training data set. At the same time, the extracted edge scenario images are preprocessed and feature vectorized with high quality, improving the data consistency and reducing the need for manual intervention.
[0055] 3. Based on the open-source large model architecture, the present invention conducts model fine-tuning training, and finally realizes intelligent vehicle scenario understanding enhanced by edge data learning. Selecting the open-source large model as a benchmark, regularly comparing its performance in basic cognition, future change prediction, and vehicle behavior analysis, and evaluating the practical application value of the new model, a complete evaluation system for existing scenario understanding tasks is constructed.
[0056] 4. Maintain the original task understanding ability: During the fine-tuning process, add a cross-entropy term to the loss function to quantify the difference between the outputs of the original model and the new model, ensuring that the learning of the new model does not deviate significantly from the performance of the original model. This not only improves the accuracy of predicting future states but also ensures the generalization ability of the model on new and old tasks.
[0057] 5. By automatically screening out challenging edge scenarios and applying advanced automation technologies for data improvement, the present invention can effectively enhance the performance of the model in dealing with rare or extreme conditions, and enhance the system's understanding ability of complex driving environments. It can be used for understanding and cognition in complex and edge scenarios, and improve the interpretability of intelligent vehicle algorithms;
[0058] 6. Combining road traffic rules, safe driving principles, and historical driving data, the system of the present invention can automatically generate safe and reasonable driving strategies that the host vehicle should adopt, ensuring the accuracy and safety of autonomous driving decisions and reducing the risk of traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0060] Figure 1 is the overall flowchart of the present invention;
[0061] Figure 2 is the ontology architecture diagram of the scenario description;
[0062] Figure 3 is the architecture diagram of the dataset construction;
[0063] Figure 4 is the schematic diagram of model comparison and verification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0065] Refer to Figures 1-4 , this embodiment provides an intelligent vehicle scenario understanding method based on enhanced edge data learning, including the following steps:
[0066] Step 1. Natural driving data collection. Based on the in-vehicle vision acquisition platform, collect the natural driving conditions in the urban area, specifically as follows:
[0067] S11. Sufficiently rich and real-world-like driving data is the key for the large model to understand real scenarios. In-vehicle data collection can best meet the authenticity requirements. Based on the Horizon vision suite, build a data collection platform for the test vehicle.
[0068] Install visual sensors provided by Horizon on the test vehicle, including high-definition cameras, stereo cameras, and surround-view cameras, to ensure that these sensors cover the full range of perspectives around the vehicle and can capture rich visual information. Configure an in-vehicle computer to process and store the collected data in real time.
[0069] S12. Capture diverse driving scenarios in real time: Through the camera devices installed on the intelligent vehicle, record the visual information during the vehicle's driving in the urban area in real time, including but not limited to various complex road traffic elements such as road environment, traffic signs, pedestrians, other vehicles, weather conditions, and road surface conditions.
[0070] S13. When selecting the vehicle driving area, ensure that the data covers a variety of driving scenarios, including but not limited to normal driving, turning, lane changing, intersection passing, obstacle avoidance, coping with emergencies (such as pedestrians suddenly crossing the road), and special weather conditions (rain, snow, fog), etc., to reflect the complexity of real-world driving.
[0071] Study and determine various driving scenarios that intelligent vehicles may face, including but not limited to normal urban streets, highways, rural roads, complex intersections, narrow lanes, curves, slopes, etc. Analyze the dynamic behaviors of various traffic participants, including the behavior patterns of pedestrians, bicycles, motorcycles, and other vehicles. At the same time, consider different climate and light conditions, including scenarios such as rain, snow, fog, night, dawn and dusk, and strong sunlight reflection. Thus, design a route loop driving including typical scenarios to ensure the sufficiency and diversity of data collection.
[0072] Step 2. Scene structure representation and semantic label processing. Based on the scene ontology model, construct a scene structured template; based on the YOLO vision processing algorithm, implement the scene semantic label classification function; train a graph neural network to fill and transform the scene images into basic elements in the template, specifically as follows:
[0073] S21. Construction of the scene structured template. Based on ontology, abstract and summarize the complex driving scenarios into a series of structured templates, as Figure 2 shown, specifically including a term layer, an attribute layer, an association layer, etc. By annotating according to this template, the spatial relationships and interaction rules between different elements can be defined, providing a standardized scene framework for subsequent intelligent analysis.
[0074] The term layer defines the basic element types and basic states in the scene; the attribute layer expands and represents the basic states; the association layer describes the association relationships between elements and the behavior specification constraints of each element. Through such an ontology, the test scenarios can be described in detail.
[0075] S22. Scene semantic label classification. Using the visual processing algorithm YOLO, perform in-depth analysis on the images obtained in the natural driving data collection stage.
[0076] Use the mature and general YOLO-v8 architecture to process the collected data. Extract image features through a Convolutional Neural Network (CNN), and use the anchor box mechanism (a preset set of bounding boxes with different sizes and aspect ratios that cover all positions in the input image) to predict the position and class probability of the bounding boxes. Extract the terms and basic attributes in the ontology architecture diagram, quickly locate and identify each object in the image, assign corresponding semantic labels to each object, and output information such as the class probability, center point coordinates (x, y), width (w), and height (h) of each type of bounding box, and convert it into a bounding box representation. Specifically, for an input image , the model outputs a series of bounding box predictions , where each bounding box contains the following information: , where represents the probability distribution of the class, is the probability distribution of the bounding box parameters, including the center point coordinates and the width and height .
[0077] S23. Image filling and transformation. Map the scene elements with semantic labels to a preset scene structured template. Based on the graph neural network, transform the originally unstructured scene image into a structured representation with clear semantic meanings and positional relationships, and convert the natural driving data into a form that can be efficiently utilized by the large model, facilitating in-depth understanding and learning by the large model.
[0078] Specifically, perform according to the following steps:
[0079] ① Treat each detected object as a node , where each node corresponds to a bounding box , and the node attributes include but are not limited to the object category, position, and size; add edges to express the spatial or semantic relationships between objects;
[0080] ② Create a graph neural network model Graph Convolutional Network (GCN) to capture the interactions between nodes and propagate context information. Define a node embedding layer to encode the original features (the bounding information output by YOLO) of each object into a high-dimensional vector representation, that is, the initial feature representation of the node , , where Represents a multi - layer perceptron in a neural network. Design a graph convolutional layer GCL to update node representations through a message - passing mechanism:
[0081] Where, is the layer index, represents the set of neighbors of node , and are the weight matrix and bias vector respectively, is the activation function. and are node features.
[0082] Is suitable for structured representation learning, enabling the network to learn complex relationships between nodes;
[0083] ③ Train the graph neural network model so that it can learn and infer deeper semantic relationships and structured information from the graph structure. After training, input the YOLO detection results into the graph neural network model, and the model will output updated node representations, which can be mapped to the concepts and their relationships in the ontology framework to form a structured representation.
[0084] Step 3: Edge scenario extraction and automatic annotation correction. Develop and apply specific metrics to automatically screen out challenging edge scenarios. For cases where the classification accuracy is lower than the set threshold, introduce a secondary review to improve data quality. Specifically as follows:
[0085] S31. Definition of edge scenarios: There is no clear definition of edge scenarios in the academic community. Considering that edge scenarios pose challenges to existing intelligent algorithms, the present invention defines edge scenarios as those situations that pose challenges to current intelligent algorithms, specifically including the following two parts:
[0086] ① Uncertain recognition: The object detection confidence output by the YOLO algorithm is lower than the set threshold (set to 0.7), indicating that the model is uncertain about whether the detected object belongs to a certain category.
[0087] ② Poor structure: After being processed by the graph neural network (GNN), the node classification accuracy fails to reach the predetermined standard (set to 90%), meaning that the elements in the image are not correctly classified and associated.
[0088] Specifically, the quantization metrics are designed as follows:
[0089] ① Confidence of object recognition: Reflects the confidence of the model that the object within the bounding box is indeed the target object rather than the background, and the calculation formula is the product of the class probability p(c) and the intersection - over - union (IoU) of the predicted box and the ground - truth box.
[0090] Specifically, the YOLO class confidence reflects the confidence level of the model that the object detected within the current bounding box is indeed the target object rather than the background. It is usually the product of the probability that the target box contains an object and the accuracy of the bounding box position. The calculation formula is as follows:
[0091] In the formula, is the class probability; is the intersection over union (IoU) between the predicted box and the ground truth box;
[0092] ② Scene structuring accuracy: Measures the performance of the GNN in correctly classifying each region or node in the image into predefined categories. It is evaluated by comparing the predicted labels with the known ground truth labels, and a confusion matrix is used to count TP, FP, and FN to calculate the accuracy.
[0093] Specifically, the graph neural network is used to structurally process the image, dividing the image into multiple nodes or regions, and each node has a predicted class label. For each node, its predicted class label is compared with the manually annotated ground truth label. If they are the same, it is counted as a correct classification. Calculate the total number of correct classifications, denoted as TP (True Positive), and at the same time count the number of classification errors, denoted as FP (False Positive, i.e., nodes misjudged as a certain category) and FN (False Negative, i.e., missed nodes). The accuracy calculation formula is as follows:
[0094] ;
[0095] S32. Automatic extraction of edge scenes: Through the above-mentioned scene edge judgment index system, automatically screen out those edge driving scene images that cannot be accurately recognized or have poor structuring effects. These edge scenes may include rare road conditions, special traffic events, views under extreme weather conditions, etc.
[0096] S33. Automatic data improvement: For the automatically extracted edge scene images, due to the poor performance of the object detection algorithm and the scene structuring algorithm, it is difficult to generate accurate and complete scene structuring data. Therefore, the system introduces automatic annotation technology to improve these data to ensure that the model can learn and adapt to complex scenes more effectively. Through weak supervision learning methods, secondary annotation classification is performed on various types of objects in the image. Through the iterative learning process, the system can gradually improve the understanding accuracy of the elements in complex scenes, ensure that all elements of each frame of the image are marked, and thus improve the knowledge base of the model when processing such complex scenes;
[0097] Step 4: Scenario Information Prediction Matching and Dataset Construction. Match future possible changes according to the scenario semantic tags, and mark the actions that the vehicle should take under the current situation. Perform preprocessing on the original images and encode them into vector format for constructing the training dataset (scenario vector data as input, structure template representation, future change data, and the main vehicle's response measures as output). Specifically as follows:
[0098] S41. Scenario Information Prediction Matching. Analyze the original collected video or image sequence based on the scenario semantic tags, extract the state changes of the scenario within a period of time, including the motion characteristics of elements such as vehicles, pedestrians, and traffic signs, and possible scenario changes, thereby constructing scenario prediction information.
[0099] S42. Main Vehicle Behavior Annotation. For the current scenario, combine road traffic rules and safe driving principles, and convert the above rules into prompt words through a pre-trained large language model, so as to automatically generate the driving strategies or actions that the main vehicle should take in a specific scenario, specifically including basic tasks such as lane change, deceleration, and parking waiting. In addition, correct the generated strategies based on the above rules and delete behaviors that do not conform to the rules.
[0100] S43. Image Preprocessing. Perform a series of preprocessing operations on the collected original image data, specifically including four steps: color space conversion, image cropping, normalization (adjusting pixel values to between 0 and 1), and size scaling (unifying the image size required for input to the neural network), to improve the efficiency of model training and inference.
[0101] S44. Feature Vectorization. Convert the preprocessed image into a numerical feature vector that can be understood by the computer, and use a convolutional neural network (CNN) to extract image features, mapping them from a high-dimensional pixel space to a low-dimensional feature space.
[0102] S45. Dataset Construction. Combine the above processing results to construct a complete dataset containing scenario information. Each sample includes not only the image data that has been preprocessed and encoded as a feature vector, but also the corresponding scenario semantic tags, predicted future state change data, and the ideal response measure tags for the main vehicle in the current scenario.
[0103] Step 5: Large Model Architecture Design and Structure Fine-tuning. Use the mature qwen-vl architecture and fine-tune it for the edge scenario dataset so that the model can learn to extract features from images and make accurate future predictions. To maintain the generalization of the trained model, add the cross-entropy of the outputs of the two models to the loss function to quantify the difference between the outputs of the original model and the new model, ensuring that the learning of the new model does not deviate significantly from the performance of the original model. Regularly verify the model performance and adjust the training parameters to ensure the generalization ability of the model until a satisfactory performance level is achieved. Specifically as follows:
[0104] S51. Model Architecture Selection and Initialization: Select an advanced and mature open-source large model architecture such as qwen-vl. Based on its powerful generalization ability and rich parameters, lay a solid foundation for subsequent scene understanding and prediction tasks. Initialize the new model based on the pre-trained qwen-vl weights.
[0105] S52. Application of the Dataset: Use the dataset constructed in Step 4 that includes the prediction of edge scene information for training to enhance the model's understanding of complex and rare driving conditions.
[0106] Utilize the Lora (Low-Rank Adaptation) method to only adjust the part of the model used for handling new tasks, especially the low-rank decomposition matrix in the fully connected layer. This method allows the network to focus on learning features related to new tasks without changing the original pre-trained weights. Adjust the dimension of the input layer according to the format of the image feature vector to be consistent with the feature vector in the aforementioned preprocessing stage. To maintain the generalization of the model, add a cross-entropy term to the standard loss function to quantify the difference between the outputs of the original model and the new model. This ensures that the learning of the new model does not deviate significantly from the performance of the original model while improving the accuracy of future state prediction.
[0107] S53. Performance Monitoring and Adjustment: During the entire training cycle, regularly (every two rounds of training) evaluate the performance of the model on an independent validation set, and pay attention to the performance metrics of the model on specific tasks (such as scene understanding accuracy, future state prediction error, and rationality score of the host vehicle's response measures). Dynamically adjust the training strategy according to the evaluation results. Use the cosine annealing learning rate scheduler to automatically adjust the learning rate based on the performance on the validation set. The learning rate gradually decreases from the initial value to the minimum value and then increases again, forming a periodic change. According to the overfitting situation, appropriately adjust the L2 regularization coefficient to prevent the model from overfitting the training data. The adjustment process gradually increases by 10% until the overfitting phenomenon disappears.
[0108] Among them, the L2 regularization coefficient is the decay of the loss weight, and the calculation method of the loss function is as follows:
[0109] In the formula, is the original loss function, is the L2 regularization coefficient, is the weight coefficient of the model. When is too large, the penalty for large weights will be heavier, resulting in a simpler model and thus underfitting. When is too large, the restriction on the weights is small, and the model will become too complex and thus overfit; therefore, it needs to be gradually increased according to the steps described above.
[0110] Step 6: Scene understanding application and model comparison evaluation. Select an open-source typical large image recognition model to analyze and understand the collected scene data, compare the analysis results of the model in terms of basic scene cognition, future scene changes, and expected vehicle behaviors, and evaluate the application effect of the model as follows:
[0111] S61. Model application test: Apply the qwen-vl model fine-tuned and trained in Step 5 and other open-source typical large image recognition models, such as EfficientNet, ResNet, etc., to the edge scene dataset constructed and processed in Steps 1 to 4 for scene understanding analysis.
[0112] S62. Scene cognition evaluation: Compare the performance of each model in basic scene recognition, and judge whether each model can accurately identify elements such as vehicles, pedestrians, traffic signs, and road conditions in the scene, and whether it can grasp the overall layout and structural characteristics of the scene. The specific index is designed as: perception category accuracy.
[0113] S63. Future change prediction comparison: Analyze the ability of each model to predict future changes in the scene, such as predicting the motion states of vehicles and pedestrians, and predicting the change trends of traffic signals, and compare the accuracies of different models in this regard. The specific indexes are designed as: moving object prediction accuracy and collision risk assessment accuracy.
[0114] S64. Vehicle behavior simulation evaluation: For each scene, evaluate the effect of each model in simulating the reasonable behaviors that an autonomous vehicle should take, and reflect whether the decision-making responses (such as lane change, braking, acceleration, etc.) in complex road conditions comply with traffic regulations and safe driving guidelines. The specific index is designed as: expected behavior rationality.
[0115] S65. Comprehensive model performance evaluation: Summarize the above evaluation results, compare the advantages and disadvantages of different models in edge scene understanding and prediction tasks from multiple dimensions, find the most suitable large model for the intelligent vehicle application scenario, and at the same time identify the deficiencies of the model to provide a direction for further optimizing the model performance. Set the weights of each index to 0.4, 0.3, and 0.3, and perform weighted averaging on the scores of each index.
[0116] In summary, the present invention can be used for understanding and cognition in complex and edge scenes, and improving the interpretability of intelligent vehicle algorithms.
[0117] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A smart car scene understanding method based on edge data learning enhancement, characterized in that: The following steps are involved: Step 1: Based on the real vehicle visual acquisition platform, collect natural driving conditions in the urban area; Step 2: Build a structured template based on the scene ontology; use the YOLO visual processing algorithm to perform target detection and semantic label classification; use the graph neural network to fill and map the scene image into the predefined scene template to complete the conversion from image to structured data; Step 3: Develop and apply indicators to automatically screen out edge scenarios; for cases where the classification accuracy is lower than the set threshold, introduce a secondary review; Step 4: Match future changes based on scene semantic labels and mark the actions that the vehicle should take under the current situation; pre-process the original image and encode it into a vector format for building a training data set; the scene vector data is used as input, and the structure template representation, future change data, and the main vehicle's response measures are used as output; Step 5: Use the mature Qwen-VL architecture and fine-tune it for edge scene datasets so that the model can learn to extract features from images and make accurate future predictions; add the cross entropy of the original model and the fine-tuned new model output to the loss function to quantify the difference between the original model and the new model output; regularly verify the model performance and adjust the training parameters until a satisfactory performance level is achieved; Step 6: Select a large open source model as a benchmark, compare its performance in basic cognition, future change prediction, and vehicle behavior analysis, and evaluate the practical application value of the model; analyze the performance difference between the new model and the original model on the same test set to confirm whether the new model successfully retains the original task understanding ability; The specific method of step 4 is as follows: S41, predicting and matching scene information; analyzing the original acquired video or image sequence based on the scene semantic label, extracting the state change of the scene in continuous time, and constructing scene prediction information; S42, labeling the main vehicle behavior; for the current scene, combining road traffic rules and safe driving principles, the pre-trained large language model is used to convert the rules into prompt words, thereby automatically generating the driving strategy or action that the main vehicle should take in the scene; in addition, the generated strategy is modified based on the rules, and the behavior that does not comply with the rules is deleted; S43, preprocessing the image, performing a series of preprocessing operations on the collected original image data, specifically including four steps of color space conversion, image cropping, normalization, and size scaling; S44, feature vectorization, converting the preprocessed image into a numerical feature vector that can be understood by the computer, using a convolutional neural network to extract image features, and mapping image features from a high-dimensional pixel space to a low-dimensional feature space; S45, constructing a data set, combining the processing results, and constructing a complete data set containing scene information; each sample includes image data that has been preprocessed and encoded into a feature vector and the corresponding scene semantic label, predicted future state change data, and an ideal response measure label that the main vehicle should take in the current scene; The specific method of step five is as follows: S51. Select and initialize the model architecture: select the qwen-vl open source large model architecture; initialize the new model based on the pre-trained qwen-vl weights; S52, application data set: use the data set containing edge scene information prediction constructed in step 4 for training; S53, Performance monitoring and adjustment: During the entire training cycle, regularly evaluate the performance of the model on an independent validation set, and pay attention to the performance indicators of the model on the data set task; dynamically adjust the training strategy based on the evaluation results, use the cosine annealing learning rate scheduler, and automatically adjust the learning rate based on the performance on the validation set. The learning rate gradually decreases from the initial value to the minimum value, and then increases again, forming a periodic change; according to the overfitting situation, adjust the L2 regularization coefficient to prevent the model from overfitting the training data. The adjustment process gradually increases by 10% until the overfitting phenomenon disappears; Among them, the L2 regularization coefficient is the attenuation of the loss weight, and the loss function is calculated as follows: Where OriginalLoss is the original loss function; λ is the L2 regularization coefficient; ωi is the weight coefficient of the model.
2. According to claim 1, a smart car scene understanding method based on edge data learning enhancement is characterized in that: The specific method of step one is as follows: S11. Build a test vehicle data acquisition platform based on the Horizon Vision Kit; S12. Real-time capture of diverse driving scenarios: Through the camera equipment installed on the smart car, the visual information of the vehicle driving in the urban area is recorded in real time; S13. When selecting the vehicle driving area, ensure that the data covers a variety of driving scenarios.
3. According to claim 1, a smart car scene understanding method based on edge data learning enhancement is characterized in that: The specific method of step 2 is as follows: S21. Construct a scenario structured template. Based on ontology, the driving scenario is abstracted into a series of structured templates, including terminology layer, attribute layer and association layer. The spatial relationship and interaction rules between different elements are defined according to this template. S22, classify the scene semantic labels and use the visual processing algorithm YOLO to perform in-depth analysis on the images obtained during the natural driving data collection phase; S23. Fill and transform images, map scene elements with semantic labels to preset scene structured templates, and transform the originally unstructured scene images into structured representations with clear semantic meanings and positional relationships based on graph neural networks, thus transforming natural driving data into a form that can be efficiently used by large models.
4. The method for intelligent vehicle scene understanding based on edge data learning enhancement according to claim 3 is characterized in that: The specific method of S22 is as follows: use the YOLO-v8 architecture to process the collected data, extract the terms and basic attributes in the ontology architecture diagram, locate and identify each target in the image, assign a corresponding semantic label to each target, and output the category probability, center point coordinates, width, and height of each type of bounding box, and convert them into a bounding box representation.
5. The method for intelligent vehicle scene understanding based on edge data learning enhancement according to claim 3 is characterized in that: The specific method of S23 is as follows: S231, taking each detected object as a node, and adding edges to express the spatial relationship or semantic relationship between the objects; S232. Create a graph neural network model to capture the interactions between nodes and propagate contextual information; define a node embedding layer to encode the original features of each object, i.e., the boundary information output by YOLO, into a high-dimensional vector representation; design a graph convolutional layer and a message passing mechanism suitable for structured representation learning, so that the network can learn the complex relationships between nodes; S233, training the graph neural network model to learn and infer deeper semantic relationships and structured information from the graph structure; After training, the YOLO detection results are input into the graph neural network model, and the model will output updated node representations that can be mapped to concepts and relationships in the ontology framework to form a structured representation.
6. The method for intelligent vehicle scene understanding based on edge data learning enhancement according to claim 1 is characterized in that: The specific method of step three is as follows: S31. Define edge scenarios: Edge scenarios are defined as those situations that pose challenges to current intelligent algorithms, specifically including the following two parts: ① Unclear identification: When the target detection confidence output by the YOLO algorithm is lower than the set threshold, it indicates that the model is uncertain about whether the detected object belongs to a certain category; ②Poor structuring: After the graph neural network is processed, when the node classification accuracy fails to reach the predetermined standard, it means that the elements in the image are not correctly classified and associated; S32, automatic extraction of edge scenes: Through the scene edge judgment index system, the edge driving scene images that cannot be accurately identified or are poorly structured are automatically screened out. Edge driving scenes include rare road conditions, poorly structured traffic events, and views under extreme weather conditions; S33. Automated data improvement: For automatically extracted edge scene images, the system introduces automated labeling technology to improve the data. Through weakly supervised learning methods, various types of targets in the image are secondary labeled and classified. Through an iterative learning process, it ensures that all elements of each frame of the image are labeled, thereby improving the model's knowledge base when dealing with such complex scenes.
7. The method for intelligent vehicle scene understanding based on edge data learning enhancement according to claim 1 is characterized in that: The specific method of step six is as follows: S61. Apply the test model: apply the Qwen-VL model obtained through fine-tuning training in step 5 and other open source typical image recognition large models to the edge scene datasets constructed and processed in steps 1 to 4, respectively, to perform scene understanding analysis; S62, cognitive evaluation scenario: compare the performance of each model in basic scene recognition, and judge whether each model can accurately identify vehicles, pedestrians, traffic signs, road conditions in the scene, and whether it can grasp the overall layout and structural characteristics of the scene; the specific indicators are designed as: perception category accuracy; S63, predict and compare future changes: analyze the ability of each model to predict future changes in the scene and compare the accuracy of different models; the specific indicators are designed as: prediction accuracy of moving subjects and collision risk assessment accuracy; S64, Simulate and evaluate vehicle behavior: For each scenario, evaluate the effectiveness of each model in simulating the reasonable behavior that the autonomous vehicle should take, reflecting whether the decision-making response under complex road conditions complies with traffic regulations and safe driving guidelines; the specific indicators are designed as follows: reasonableness of expected behavior; S65. Comprehensively evaluate model performance: Summarize the evaluation results, compare the advantages and disadvantages of different models in edge scene understanding and prediction tasks from multiple dimensions, find the large model that is most suitable for smart car application scenarios, set the weight of each indicator to 0.4, 0.3, and 0.3, and take a weighted average of the scores of each indicator.
Citation Information
Patent Citations
Multi-scene reflective vest wearing identification model creation method and related components
CN113158993A
Haze scene driving vision enhancement and target detection method
CN117974497A