Model training and usage methods, devices, media, and equipment for driving scenario analysis

By encoding elements in the autonomous driving scenario into nodes and representing connection relationships, and combining the analysis model to learn scene characteristics, the problems of insufficient information and sparse information are solved, and the accuracy of autonomous driving scenario analysis and computing resource utilization are improved.

CN116152601BActive Publication Date: 2025-06-17JIUZHI (SUZHOU) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310144481.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-06-17
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

In the analysis of autonomous driving scenarios, the prior art has problems such as too little information, sparse information, and the inability to express interaction behavior.

Method used

An encoded graph is used to represent the autonomous driving scenario by encoding traffic participants and static map elements into nodes and representing the connection relationship between each node through connections. The coded graph is processed using an analytical model to learn scene characteristics, node characteristics and connection relationship characteristics.

Benefits of technology

It effectively avoids the problems of missing front-view camera information and sparse grid image information, enhances the ability to express scene characteristics, improves the accuracy of autonomous driving scenario classification, and optimizes the utilization of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152601B_ABST
    Figure CN116152601B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, medium, and device for model training and use in driving scenario analysis, belonging to the technical field of machine learning. The method includes: obtaining a training set, where each training sample includes an encoded map of an autonomous driving scenario and first node features, the nodes in the encoded map represent traffic participants and static map elements, and the connections represent the connection relationships between the nodes; for each training sample, using the encoding model in the analysis model to process the encoded map to obtain scene features; using the decoding model in the analysis model to process the scene features to obtain the connection relationships between the nodes and second node features of each node; adjusting the loss function according to the encoded map, first node features, connection relationships, and second node features; and training the analysis model according to the loss function. The present application uses an encoded map to represent an autonomous driving scenario, which can avoid information loss and information sparsity and can express interaction behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and particularly to a method, device, medium, and equipment for training and using a model for driving scenario analysis. Background Art

[0002] With the development of autonomous driving technology, machine learning and deep learning technologies are increasingly widely used in autonomous driving. These data-driven technologies and data together constitute the data closed-loop of autonomous driving. In the data closed-loop, we often need to analyze autonomous driving scenarios. Due to the diversity of road networks and the complexity of the behaviors of traffic participants, how to use scene features to represent a highly interactive scene is a difficult problem in autonomous driving scene analysis. It is also the key point restricting the collection of long-tail problems in autonomous driving interactivity.

[0003] In related technologies, the following two methods are adopted to extract scene features:

[0004] (1) Feature extraction and classification are performed based on the data captured by the front-view camera in one frame or a sequence. Among them, the method of feature extraction is to use deep neural networks such as CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network) to process the front-view image or video information to obtain high-dimensional scene features. The image obtained by the front-view camera is as Figure 1 shown.

[0005] (2) Road features and traffic participants are projected onto a grid map in the BEV (Bird's eye view) perspective, and the scene features are represented based on the grid map. The grid map in the BEV is as Figure 2 shown.

[0006] For the first method of scene feature extraction, the data captured by the front-view camera is from one perspective and cannot represent traffic participants in other perspectives; moreover, it cannot express the road structure and lacks important environmental information, that is, its main disadvantage is that the amount of information is too small. For the second method of scene feature extraction, it makes up for the problem of insufficient perspective, but most of the image area of the grid map has no road or traffic participant information, so the available information is very sparse, and one frame of the grid map cannot effectively represent the interaction behaviors between traffic participants, that is, its main disadvantages are sparse information and inability to express interaction behaviors, etc. Summary of the Invention

[0007] The present application provides a method, apparatus, medium, and device for model training and use in driving scenario analysis, which are used to solve the problems of too little information, sparse information, and inability to express interaction behaviors. The technical solutions are as follows:

[0008] On the one hand, a method for model training in driving scenario analysis is provided. The method includes:

[0009] Obtain a training set, where each training sample in the training set includes an encoded map of an autonomous driving scenario and multiple first node features. The autonomous driving scenario is a scenario composed of the perception results of an autonomous vehicle over a period of time. The nodes in the encoded map represent traffic participants and static map elements in the perception results, the connections in the encoded map represent the connection relationships between the respective nodes, and the first node features represent the features of the nodes;

[0010] Create an analysis model;

[0011] For each training sample, use the encoding model in the analysis model to process the encoded map to obtain the scenario features of the autonomous driving scenario; use the decoding model in the analysis model to process the scenario features to obtain the connection relationships between the respective nodes and the second node features of the respective nodes; adjust the loss function of the analysis model according to the encoded map, the first node features, the connection relationships, and the second node features; train the analysis model according to the loss function.

[0012] In a possible implementation, the step of using the encoding model in the analysis model to process the encoded map to obtain the scenario features of the autonomous driving scenario includes:

[0013] When the encoding model includes two graph convolutional neural networks, use the two graph convolutional neural networks to perform feature extraction on the encoded map to obtain the scenario features of the autonomous driving scenario.

[0014] In a possible implementation, the step of using the decoding model in the analysis model to process the scenario features to obtain the connection relationships between the respective nodes and the second node features of the respective nodes includes:

[0015] When the decoding model includes a first decoding branch and a second decoding branch, use the multi-layer perceptron in the first decoding branch to process the scenario features to obtain the connection relationships between the respective nodes; use the multi-layer perceptron and convolutional neural network in the second decoding branch to process the scenario features to obtain the second node features of the respective nodes.

[0016] In a possible implementation, adjusting the loss function of the analysis model according to the encoded graph, the first node features, the connection relationship, and the second node features includes:

[0017] Calculating the matching degree between the connection relationship and the connection relationship in the encoded graph, and calculating the matching degree between the first node features and the second node features;

[0018] Adjusting the loss function of the analysis model according to the matching degree.

[0019] In a possible implementation, the method further includes:

[0020] Obtaining the perception result;

[0021] Extracting the third node features of the traffic participants from the perception result, and extracting the fourth node features of the static map elements from the perception result or from the high-precision map;

[0022] Processing the third node features by using a long short-term memory network to obtain the first node features of the traffic participants;

[0023] Processing the fourth node features by using a convolutional neural network to obtain the first node features of the static map elements.

[0024] On the one hand, a method for using a model for driving scenario analysis is provided, and the method includes:

[0025] Obtaining the perception results of an autonomous vehicle over a period of time;

[0026] Generating an encoded graph according to the perception result, where the nodes in the encoded graph represent the traffic participants and static map elements in the perception result, and the connections in the encoded graph represent the connection relationships between the respective nodes;

[0027] Processing the encoded graph by using the encoding model in the analysis model to obtain the scene features of each node, and the scene features are used to analyze the autonomous driving scenario.

[0028] In a possible implementation, the method further includes:

[0029] When the decoding model in the analysis model includes a first decoding branch, using the multi-layer perceptron in the first decoding branch to process the scene features to obtain the connection relationships between the respective nodes.

[0030] In a possible implementation, the method further includes:

[0031] When the decoding model in the analysis model includes a second decoding branch, use the multi-layer perceptron and convolutional neural network in the second decoding branch to process the scene features to obtain second node features of each node, and the second node features are used to analyze the node.

[0032] On the one hand, a model training device for driving scene analysis is provided, and the device includes:

[0033] An acquisition module for acquiring a training set, where each training sample in the training set includes an encoded map of an autonomous driving scene and multiple first node features. The autonomous driving scene is a scene composed of the perception results of an autonomous driving vehicle over a period of time. The nodes in the encoded map represent traffic participants and static map elements in the perception results, the connections in the encoded map represent the connection relationships between the respective nodes, and the first node features represent the features of the nodes.

[0034] A creation module for creating an analysis model.

[0035] A training module for, for each training sample, using the encoding model in the analysis model to process the encoded map to obtain the scene features of the autonomous driving scene; using the decoding model in the analysis model to process the scene features to obtain the connection relationships between the respective nodes and the second node features of each node; adjusting the loss function of the analysis model according to the encoded map, the first node features, the connection relationships, and the second node features; and training the analysis model according to the loss function.

[0036] On the one hand, a model usage device for driving scene analysis is provided, and the device includes:

[0037] An acquisition module for acquiring the perception results of an autonomous driving vehicle over a period of time.

[0038] A generation module for generating an encoded map according to the perception results, where the nodes in the encoded map represent traffic participants and static map elements in the perception results, and the connections in the encoded map represent the connection relationships between the respective nodes.

[0039] An analysis module for using the encoding model in the analysis model to process the encoded map to obtain the scene features of each node, and the scene features are used to analyze the autonomous driving scene.

[0040] On the one hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the model training method for driving scenario analysis as described above; or, the at least one instruction is loaded and executed by a processor to implement the model usage method for driving scenario analysis as described above.

[0041] On the one hand, a computer device is provided, the computer device includes a processor and a memory, and at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the model training method for driving scenario analysis as described above; or, the instruction is loaded and executed by the processor to implement the model usage method for driving scenario analysis as described above.

[0042] The beneficial effects of the technical solution provided by this application at least include:

[0043] By encoding traffic participants and static map elements in an autonomous driving scenario as nodes, and then representing the connection relationships of each node through connections, a coding graph can be used to represent an autonomous driving scenario, which can not only avoid the problem of information loss when using a front view camera, but also avoid the problem of sparse information caused by using a raster map and the inability to express interaction behaviors, and can also avoid the waste of computing resources caused by sparse information. Moreover, an analysis model can be used to learn the scene features of the coding graph, which can enhance the expression of scene features and improve the accuracy of autonomous driving scenario classification; it can also learn the nodes and connection relationships, can perform a complete high-dimensional abstraction of the features of the coding graph, and uses the structure of the coding graph to describe time and space information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 is a schematic diagram of an image captured by a front view camera;

[0046] Figure 2 is a schematic diagram of a raster map under BEV;

[0047] Figure 3 is a flowchart of the model training method for driving scenario analysis provided by an embodiment of the present application;

[0048] Figure 4 is a schematic diagram of generating the first node feature;

[0049] Figure 5 It is a schematic diagram of an encoded image;

[0050] Figure 6 It is a schematic structural diagram of an analysis model;

[0051] Figure 7 It is a flowchart of a method for using a model for driving scenario analysis provided by another embodiment of the present application;

[0052] Figure 8 It is a block diagram of the structure of a model training device for driving scenario analysis provided by still another embodiment of the present application;

[0053] Figure 9 It is a block diagram of the structure of a model using device for driving scenario analysis provided by still another embodiment of the present application. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0055] Autonomous vehicles are usually equipped with sensing devices such as cameras and lidar. During driving, the sensing devices sense the environment around the autonomous vehicle and transmit the sensing results to the computer device in the autonomous vehicle in real time. The computer device processes the sensing results and analyzes the autonomous driving scenario based on the processing results. Among them, the sensing result corresponds to the sensing device. For example, when the sensing device is a camera, the sensing result can be a video stream; when the sensing device is a lidar, the sensing result can be a point cloud stream.

[0056] In this embodiment, the scenario formed by the sensing results within a period of time is called an autonomous driving scenario. For example, when the autonomous vehicle drives to an intersection, the computer device transmits the sensing results captured during driving to the computer device. After the computer device processes the sensing results, it classifies the autonomous driving scenario. The classification here can include determining to stop driving according to the traffic lights, or determining to meet the right-turn condition according to the traffic conditions at the intersection, etc.

[0057] In this embodiment, the computer device can first convert the sensing result into a structured encoded image, and then use the analysis model to analyze the encoded image. The following describes the training process of the analysis model.

[0058] Please refer to Figure 3 , which shows a flowchart of a method for training a model for driving scenario analysis provided by an embodiment of the present application. The method for training a model for driving scenario analysis can be applied to a computer device. The method for training a model for driving scenario analysis can include:

[0059] Step 301: Obtain a training set. Each training sample in the training set includes an encoded map of an autonomous driving scenario and multiple first node features. The autonomous driving scenario is a scenario composed of the perception results of an autonomous vehicle over a period of time. The nodes in the encoded map represent traffic participants and static map elements in the perception results, the connections in the encoded map represent the connection relationships between the respective nodes, and the first node features represent the features of the nodes.

[0060] Before training the analysis model, the computer device needs to first generate training samples, which include an encoded map and multiple first node features. Among them, each sample is generated based on the perception results over a period of time. Specifically, generating training samples may include the following steps:

[0061] (1) Obtain the perception results.

[0062] Among them, the perception results are obtained by the autonomous vehicle perceiving the surrounding environment for a period of time during driving.

[0063] (2) Extract the third node features of traffic participants from the perception results, and extract the fourth node features of static map elements from the perception results or from the high-precision map.

[0064] Taking the perception results as a video as an example, the computer device samples an image sequence from the video stream according to the sampling frequency, then identifies the objects in the image sequence, and extracts the node features of the identified traffic participants. In this embodiment, this node feature is called the third node feature. The computer device then extracts static map elements from the high-precision map and extracts the node features of the static map elements. In this embodiment, this node feature is called the fourth node feature. Or, the computer device extracts an image sequence from the video stream according to the sampling frequency, then identifies the objects in the image sequence, and extracts the node features of the identified traffic participants and static map elements. In this embodiment, the node feature of the traffic participant is called the third node feature, and the node feature of the static map element is called the fourth node feature. Among them, traffic participants can be moving or stationary vehicles or pedestrians, such as a moving or stopped autonomous vehicle, etc. Static map elements are elements in the traffic road environment, such as lanes, traffic lights, signs, etc.

[0065] When extracting the third node features, continuous time information needs to be used. Assume that the node feature of a traffic participant agent at each moment is f t = [x, y, v x , v y , heading,...], where x represents the abscissa of the autonomous vehicle, y represents the ordinate of the autonomous vehicle, v x represents the lateral driving speed of the autonomous vehicle, vy The vertical driving speed of the autonomous vehicle is represented, and heading represents the orientation of the autonomous vehicle. In this embodiment, only the above parameters are used for illustration. In actual business, more or fewer parameters can be selected to represent the node features of the autonomous vehicle. In an autonomous driving scenario, the third node feature Fagent of a traffic participant = [f t , f t-1 , …, f t-n , where t represents the time and n represents the number of image sequences.

[0066] Assume that the static map element is a lane. Then the node feature of the lane is fs = [x, y, width, …], where x represents the abscissa of the sampling point in the lane, y represents the ordinate of the sampling point in the lane, and width represents the lane width corresponding to the sampling point in the lane. In this embodiment, only the above parameters are used for illustration. In actual business, more or fewer parameters can be selected to represent the node features of the lane. In an autonomous driving scenario, the fourth node feature Fmap of a lane = [f s 0 , …, f s m , where m represents the number of sampling points in the lane.

[0067] (3) Process the third node feature using a Long Short-Term Memory (LSTM) network to obtain the first node feature of the traffic participant.

[0068] Please refer to Figure 4 , the third node feature Agent_0Feature of traffic participant Agent_0 is converted by LSTM to obtain the first node feature P_agent0; the third node feature Agent_1Feature of traffic participant Agent_1 is converted by LSTM to obtain the first node feature P_agent1; the third node feature Agent_2Feature of traffic participant Agent_2 is converted by LSTM to obtain the first node feature P_agent2.

[0069] (4) Process the fourth node feature using a Convolutional Neural Network (CNN) to obtain the first node feature of the static map element.

[0070] Among them, the convolutional neural network can be a one-dimensional convolutional neural network.

[0071] Please refer to Figure 4, the fourth node feature of the static map element Map_1 is transformed by 1d_CNN to obtain the first node feature P_map1; the fourth node feature of the static map element Map_2 is transformed by 1d_CNN to obtain the first node feature P_map2.

[0072] After obtaining the first node feature of the node, the computer device also needs to create the connection relationship between nodes to obtain the encoded graph, as Figure 5 shown. Specifically, the connection relationship between nodes is divided into the following three types:

[0073] (a) Connection relationship between traffic participants: A connection relationship needs to be created between nodes that affect the driving behavior of traffic participants.

[0074] (b) Connection relationship between static map elements: A connection relationship needs to be created between adjacent lanes and lanes with successor relationships.

[0075] (c) Connection relationship between traffic participants and static map elements: A connection relationship needs to be created between traffic participants on the lane and the lane, and a connection relationship needs to be created between traffic participants not on the lane and the lane where they will be in the future.

[0076] Step 302, create an analysis model.

[0077] The analysis model includes an encoding model and a decoding model. The encoding model includes two connected Graph Convolutional Networks (GCN); the decoding model includes a first decoding branch and a second decoding branch. The first decoding branch includes a Multilayer Perceptron (MLP), and the second decoding branch includes a connected Multilayer Perceptron and a one-dimensional convolutional neural network, as Figure 6 shown.

[0078] Step 303, for each training sample, use the encoding model in the analysis model to process the encoded graph to obtain the scene feature of the autonomous driving scene; use the decoding model in the analysis model to process the scene feature to obtain the connection relationship between each node and the second node feature of each node; adjust the loss function of the analysis model according to the encoded graph, the first node feature, the connection relationship, and the second node feature; train the analysis model according to the loss function.

[0079] Specifically, using the encoding model in the analysis model to process the encoded graph to obtain the scene feature of the autonomous driving scene includes: when the encoding model includes two graph convolutional neural networks, use the two graph convolutional neural networks to extract features from the encoded graph to obtain the scene feature (Node Latent Space) of the autonomous driving scene.

[0080] Specifically, the decoding model in the analysis model is used to process the scene features, and the connection relationships between the nodes and the second node features of each node are obtained, including: when the decoding model includes a first decoding branch and a second decoding branch, the multi-layer perceptron in the first decoding branch is used to process the scene features to obtain the connection relationships (Link prediction) between the nodes; the multi-layer perceptron and the convolutional neural network in the second decoding branch are used to process the scene features to obtain the second node features (Feature prediction) of each node.

[0081] Specifically, the loss function of the analysis model is adjusted according to the encoded graph, the first node features, the connection relationships, and the second node features, including: calculating the matching degree between the connection relationships and the connection relationships in the encoded graph, and calculating the matching degree between the first node features and the second node features; adjusting the loss function of the analysis model according to the matching degree.

[0082] Among them, the analysis model adopts an unsupervised learning method, and the obtained high-dimensional scene features can be effectively used for common scene expression tasks such as matching, classification, and regression.

[0083] In summary, the model training method for driving scene analysis provided by the embodiments of the present application encodes traffic participants and static map elements in an autonomous driving scene into nodes, and then represents the connection relationships between the nodes by connecting lines, so that an encoded graph can be used to represent an autonomous driving scene, which can not only avoid the problem of information loss when using a front view camera, but also avoid the problem of sparse information caused by using a grid map and the inability to express interaction behaviors, and can also avoid the waste of computing resources caused by sparse information. Moreover, an analysis model can be used to learn the scene features of the encoded graph, which can enhance the expression of the scene features and improve the accuracy of autonomous driving scene classification; it can also learn the nodes and connection relationships, can perform a complete high-dimensional abstraction of the features of the encoded graph, and describes the time and space information by using the structure of the encoded graph.

[0084] Please refer to Figure 7 , which shows a flowchart of a method for using a model for driving scene analysis provided by an embodiment of the present application. The method for using the model for driving scene analysis can be applied to a computer device. The method for using the model for driving scene analysis can include:

[0085] Step 701, obtain the perception results of the autonomous driving vehicle within a period of time.

[0086] Step 702, generate an encoded graph according to the perception results. The nodes in the encoded graph represent the traffic participants and static map elements in the perception results, and the connecting lines in the encoded graph represent the connection relationships between the nodes.

[0087] Among them, the computer device generating the encoded map includes the following steps:

[0088] (1) Extract the third node features of traffic participants from the perception results, and extract the fourth node features of static map elements from the perception results or from the high-precision map.

[0089] Taking the perception result as a video stream as an example, the computer device samples an image sequence from the video stream according to the sampling frequency, then identifies the objects in the image sequence, and extracts the node features of the identified traffic participants. In this embodiment, this node feature is referred to as the third node feature. The computer device then extracts the static map elements from the high-precision map and extracts the node features of the static map elements. In this embodiment, this node feature is referred to as the fourth node feature. Alternatively, the computer device extracts an image sequence from the video stream according to the sampling frequency, then identifies the objects in the image sequence, and extracts the node features of the identified traffic participants and static map elements. In this embodiment, the node feature of the traffic participant is referred to as the third node feature, and the node feature of the static map element is referred to as the fourth node feature. Among them, traffic participants can be moving or stationary vehicles or pedestrians, such as, for example, a self-driving vehicle that is driving or stopped. Static map elements are elements in the traffic road environment, such as, for example, lanes, traffic lights, signs, etc.

[0090] When extracting the third node features, continuous time information needs to be used. Assume that the node feature of a traffic participant agent at each moment is f t = [x, y, v x , v y , heading, …], where x represents the abscissa of the self-driving vehicle, y represents the ordinate of the self-driving vehicle, v x represents the lateral driving speed of the self-driving vehicle, v y represents the vertical driving speed of the self-driving vehicle, and heading represents the orientation of the self-driving vehicle. In this embodiment, only the above parameters are used for illustration. In actual business, more or fewer parameters can be selected to represent the node features of the self-driving vehicle. In a self-driving scenario, the third node feature Fagent of a traffic participant = [f t , f t-1 , …, f t-n , where t represents the moment and n represents the number of image sequences.

[0091] Assume that the static map element is a lane. The node features of the lane are fs = [x, y, width,...], where x represents the abscissa of the sampling point in the lane, y represents the ordinate of the sampling point in the lane, and width represents the lane width corresponding to the sampling point in the lane. In this embodiment, only the above parameters are used for illustration. In actual business, more or fewer parameters can be selected to represent the node features of the lane. In an autonomous driving scenario, the fourth node feature Fmap of a lane is [f s 0 , …, f s m , where m represents the number of sampling points in the lane.

[0092] (2) Use a long short-term memory network to process the third node feature to obtain the first node feature of the traffic participant.

[0093] Assume that three traffic participants are recognized from the image sequence. The third node feature Agent_0Feature of traffic participant Agent_0 is converted by LSTM to obtain the first node feature P_agent0; the third node feature Agent_1Feature of traffic participant Agent_1 is converted by LSTM to obtain the first node feature P_agent1; the third node feature Agent_2Feature of traffic participant Agent_2 is converted by LSTM to obtain the first node feature P_agent2.

[0094] (2) Use a convolutional neural network to process the fourth node feature to obtain the first node feature of the static map element.

[0095] Among them, the convolutional neural network can be a one-dimensional convolutional neural network.

[0096] Assume that two static map elements are recognized from the image sequence. The fourth node feature of static map element Map_1 is converted by 1d_CNN to obtain the first node feature P_map1; the fourth node feature of static map element Map_2 is converted by 1d_CNN to obtain the first node feature P_map2.

[0097] After obtaining the first node feature of the node, the computer device also needs to create the connection relationship between the nodes to obtain the encoded graph. Specifically, the connection relationship between the nodes is divided into the following three types:

[0098] (a) Connection relationship between traffic participants: A connection relationship needs to be created between the nodes that have an impact on the driving behavior of the traffic participants.

[0099] (b) Connection relationship between static map elements: A connection relationship needs to be created between adjacent lanes and lanes with predecessor and successor relationships.

[0100] (c) Connection relationship between traffic participants and static map elements: A connection relationship needs to be created between traffic participants on the lane and the lane, and a connection relationship needs to be created between traffic participants not on the lane and the lane where they will be in the future.

[0101] Step 703: Process the encoded graph using the encoding model in the analysis model to obtain the scene features of each node. The scene features are used to analyze the autonomous driving scene.

[0102] Specifically, when the encoding model includes two graph convolutional neural networks, use the two graph convolutional neural networks to extract features from the encoded graph to obtain the scene features of the autonomous driving scene. After obtaining the scene features, the computer device can use the scene features for common scene expression tasks such as matching, classification, and regression.

[0103] If the connection relationship between each node needs to be obtained, the computer device uses the multi-layer perceptron in the first decoding branch to process the scene features to obtain the connection relationship between each node.

[0104] If the second node features of each node need to be obtained, the computer device uses the multi-layer perceptron and convolutional neural network in the second decoding branch to process the scene features to obtain the second node features of each node. The second node features are used to analyze the nodes.

[0105] In summary, the method for using the model for driving scene analysis provided by the embodiments of the present application encodes traffic participants and static map elements in an autonomous driving scene as nodes, and then represents the connection relationship between each node by connecting lines, so that an encoded graph can be used to represent an autonomous driving scene. This can not only avoid the problem of information loss when using a front-view camera, but also avoid the problem of sparse information caused by using a raster map and the inability to express interaction behaviors. It can also avoid the waste of computing resources caused by sparse information. Moreover, an analysis model can be used to learn the scene features of the encoded graph, which can enhance the expression of the scene features and improve the accuracy of autonomous driving scene classification; it can also learn the nodes and connection relationships, can perform a complete high-dimensional abstraction of the features of the encoded graph, and describes the time and space information using the structure of the encoded graph.

[0106] Please refer to Figure 8 , which shows the structural block diagram of the model training device for driving scene analysis provided by an embodiment of the present application. The model training device for driving scene analysis can be applied to a computer device. The model training device for driving scene analysis may include:

[0107] An acquisition module 810, configured to acquire a training set. Each training sample in the training set includes an encoded map of an autonomous driving scenario and multiple first node features. The autonomous driving scenario is a scenario formed by the perception results of an autonomous vehicle over a period of time. The nodes in the encoded map represent traffic participants and static map elements in the perception results, the connections in the encoded map represent the connection relationships between the respective nodes, and the first node features represent the features of the nodes.

[0108] A creation module 820, configured to create an analysis model.

[0109] A training module 830, configured to, for each training sample, use the encoding model in the analysis model to process the encoded map to obtain the scenario features of the autonomous driving scenario; use the decoding model in the analysis model to process the scenario features to obtain the connection relationships between the respective nodes and the second node features of the respective nodes; adjust the loss function of the analysis model according to the encoded map, the first node features, the connection relationships, and the second node features; and train the analysis model according to the loss function.

[0110] In an optional embodiment, the training module 830 is further configured to:

[0111] When the encoding model includes two graph convolutional neural networks, use the two graph convolutional neural networks to perform feature extraction on the encoded map to obtain the scenario features of the autonomous driving scenario.

[0112] In an optional embodiment, the training module 830 is further configured to:

[0113] When the decoding model includes a first decoding branch and a second decoding branch, use the multi-layer perceptron in the first decoding branch to process the scenario features to obtain the connection relationships between the respective nodes; use the multi-layer perceptron and the convolutional neural network in the second decoding branch to process the scenario features to obtain the second node features of the respective nodes.

[0114] In an optional embodiment, the training module 830 is further configured to:

[0115] Calculate the matching degree between the connection relationships and the connection relationships in the encoded map, and calculate the matching degree between the first node features and the second node features;

[0116] Adjust the loss function of the analysis model according to the matching degree.

[0117] In an optional embodiment, the acquisition module 810 is further configured to:

[0118] Acquire the perception results;

[0119] Extract the third node features of traffic participants from the perception results, and extract the fourth node features of static map elements from the perception results or from a high-precision map.

[0120] Process the features of the third node using a long short-term memory network to obtain the first node features of traffic participants;

[0121] Process the features of the fourth node using a convolutional neural network to obtain the first node features of static map elements.

[0122] In summary, the model training device for driving scenario analysis provided by the embodiments of the present application encodes traffic participants and static map elements in an autonomous driving scenario as nodes, and then represents the connection relationships between each node through connections. Thus, an encoding graph can be used to represent an autonomous driving scenario, which can not only avoid the problem of information loss when using a front view camera, but also avoid the problem of sparse information caused by using a grid map and the inability to express interaction behaviors. It can also avoid the waste of computing resources caused by sparse information. Moreover, an analysis model can be used to learn the scenario features of the encoding graph, which can enhance the expression of scenario features and improve the accuracy of autonomous driving scenario classification; it can also learn the nodes and connection relationships, can perform a complete high-dimensional abstraction of the features of the encoding graph, and describe time and space information using the structure of the encoding graph.

[0123] Please refer to Figure 9 , which shows a structural block diagram of a model usage device for driving scenario analysis provided by an embodiment of the present application. The model usage device for driving scenario analysis can be applied to a computer device. The model usage device for driving scenario analysis may include:

[0124] An acquisition module 910, configured to acquire the perception results of an autonomous driving vehicle within a period of time;

[0125] A generation module 920, configured to generate an encoding graph according to the perception results. The nodes in the encoding graph represent traffic participants and static map elements in the perception results, and the connections in the encoding graph represent the connection relationships between each node;

[0126] An analysis module 930, configured to process the encoding graph using the encoding model in the analysis model to obtain the scenario features of each node, and the scenario features are used to analyze the autonomous driving scenario.

[0127] In an optional embodiment, the analysis module 930 is further configured to:

[0128] When the decoding model in the analysis model includes a first decoding branch, process the scenario features using a multi-layer perceptron in the first decoding branch to obtain the connection relationships between each node.

[0129] In an optional embodiment, the analysis module 930 is further configured to:

[0130] When the decoding model in the analysis model includes a second decoding branch, the multi-layer perceptron and convolutional neural network in the second decoding branch are used to process the scene features to obtain the second node features of each node, and the second node features are used for analyzing the nodes.

[0131] In summary, the model usage device for driving scene analysis provided by the embodiments of the present application encodes traffic participants and static map elements in an autonomous driving scene into nodes, and then represents the connection relationships of each node through connections, so that an autonomous driving scene can be represented by an encoded graph. This can not only avoid the problem of information loss when using a front-view camera, but also avoid the problem of sparse information caused by using a raster map and the inability to express interaction behaviors. It can also avoid the waste of computing resources caused by sparse information. Moreover, an analysis model can be used to learn the scene features of the encoded graph, which can enhance the expression of scene features and improve the accuracy of autonomous driving scene classification. It can also learn the nodes and connection relationships, perform a complete high-dimensional abstraction of the features of the encoded graph, and describe the time and space information using the structure of the encoded graph.

[0132] An embodiment of the present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the model training method for driving scene analysis as described above; or, the at least one instruction is loaded and executed by a processor to implement the model usage method for driving scene analysis as described above.

[0133] An embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the model training method for driving scene analysis as described above; or, the instruction is loaded and executed by the processor to implement the model usage training method for driving scene analysis as described above.

[0134] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.

[0135] The above description is not intended to limit the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the embodiments of the present application.

Claims

1. A method for training a model for driving scenario analysis, characterized in that, The method includes: Obtain a training set, where each training sample in the training set includes an encoded map of an autonomous driving scenario and multiple first node features. The autonomous driving scenario is a scenario composed of the perception results of an autonomous vehicle over a period of time. The nodes in the encoded map represent the traffic participants and static map elements in the perception results, the connections in the encoded map represent the connection relationships between the respective nodes, and the first node features represent the features of the nodes; Create an analysis model; For each training sample, use the encoding model in the analysis model to process the encoded map to obtain the scenario features of the autonomous driving scenario; use the decoding model in the analysis model to process the scenario features to obtain the connection relationships between the respective nodes and the second node features of the respective nodes; adjust the loss function of the analysis model according to the encoded map, the first node features, the connection relationships, and the second node features; train the analysis model according to the loss function; The step of using the decoding model in the analysis model to process the scenario features to obtain the connection relationships between the respective nodes and the second node features of the respective nodes includes: when the decoding model includes a first decoding branch and a second decoding branch, use the multi-layer perceptron in the first decoding branch to process the scenario features to obtain the connection relationships between the respective nodes; use the multi-layer perceptron and convolutional neural network in the second decoding branch to process the scenario features to obtain the second node features of the respective nodes.

2. The method for training a model for driving scenario analysis according to claim 1, characterized in that, The step of using the encoding model in the analysis model to process the encoded map to obtain the scenario features of the autonomous driving scenario includes: When the encoding model includes two graph convolutional neural networks, use the two graph convolutional neural networks to perform feature extraction on the encoded map to obtain the scenario features of the autonomous driving scenario.

3. The method for training a model for driving scenario analysis according to claim 1, characterized in that, The step of adjusting the loss function of the analysis model according to the encoded map, the first node features, the connection relationships, and the second node features includes: Calculate the matching degree between the connection relationships and the connection relationships in the encoded map, and calculate the matching degree between the first node features and the second node features; Adjust the loss function of the analysis model according to the matching degree.

4. The method for training a model for driving scenario analysis according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the perception results; Extract the third node features of the traffic participants from the perception results, and extract the fourth node features of the static map elements from the perception results or from a high-precision map; Use a long short-term memory network to process the third node features to obtain the first node features of the traffic participants; Use a convolutional neural network to process the fourth node features to obtain the first node features of the static map elements.

5. A method for using a model for driving scenario analysis, characterized in that, The method includes: Obtain the perception results of an autonomous vehicle over a period of time; Generate an encoded map according to the perception results, where the nodes in the encoded map represent the traffic participants and static map elements in the perception results, and the connections in the encoded map represent the connection relationships between the respective nodes; Process the encoded graph using the encoding model in the analysis model to obtain the scene features of each node, where the scene features are used to analyze the autonomous driving scenario, and the analysis model is obtained by using the model training method described in any one of claims 1 to 4; When the decoding model in the analysis model includes a first decoding branch and a second decoding branch, process the scene features using the multi-layer perceptron in the first decoding branch to obtain the connection relationships between the nodes; process the scene features using the multi-layer perceptron and the convolutional neural network in the second decoding branch to obtain the second node features of the nodes, where the second node features are used to analyze the nodes.

6. A device for training a model for driving scenario analysis, characterized in that, The device includes: An acquisition module, configured to acquire a training set, where each training sample in the training set includes an encoded graph of an autonomous driving scenario and multiple first node features. The autonomous driving scenario is a scenario composed of the perception results of an autonomous vehicle over a period of time. The nodes in the encoded graph represent the traffic participants and static map elements in the perception results, the connections in the encoded graph represent the connection relationships between the nodes, and the first node features represent the features of the nodes; A creation module, configured to create an analysis model; A training module, for each training sample, process the encoded graph using the encoding model in the analysis model to obtain the scene features of the autonomous driving scenario; process the scene features using the decoding model in the analysis model to obtain the connection relationships between the nodes and the second node features of the nodes; adjust the loss function of the analysis model according to the encoded graph, the first node features, the connection relationships, and the second node features; train the analysis model according to the loss function; The training module is further configured to: when the decoding model includes a first decoding branch and a second decoding branch, process the scene features using the multi-layer perceptron in the first decoding branch to obtain the connection relationships between the nodes; process the scene features using the multi-layer perceptron and the convolutional neural network in the second decoding branch to obtain the second node features of the nodes.

7. A device for using a model for driving scenario analysis, characterized in that, The device includes: An acquisition module, configured to acquire the perception results of an autonomous vehicle over a period of time; A generation module, configured to generate an encoded graph according to the perception results, where the nodes in the encoded graph represent the traffic participants and static map elements in the perception results, and the connections in the encoded graph represent the connection relationships between the nodes; An analysis module, configured to process the encoded graph using the encoding model in the analysis model to obtain the scene features of each node, where the scene features are used to analyze the autonomous driving scenario, and the analysis model is obtained by using the model training method described in any one of claims 1 to 4; The analysis module is further configured to, when the decoding model in the analysis model includes a first decoding branch and a second decoding branch, process the scene features by using the multi-layer perceptron in the first decoding branch to obtain the connection relationships between the respective nodes; and process the scene features by using the multi-layer perceptron and the convolutional neural network in the second decoding branch to obtain second node features of the respective nodes, where the second node features are used to analyze the nodes.

8. A computer-readable storage medium, characterized in that,At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the model training method for driving scene analysis according to any one of claims 1 to 4; or, the at least one instruction is loaded and executed by a processor to implement the model usage method for driving scene analysis according to claim 5.

9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the model training method for driving scene analysis according to any one of claims 1 to 4; or, the instruction is loaded and executed by the processor to implement the model usage method for driving scene analysis according to claim 5.

Citation Information

Patent Citations

  • Image description method and system based on two-way feature encoder

    CN113642630A

  • Automatic driving scene generation method and related model training method and device

    CN115438569A