Training Method, Device, Equipment and Storage Medium for Autonomous Driving Model
By building a visual Q&A network diagram and training an autonomous driving model, the problem of insufficient generalization and interactivity of the autonomous driving model in complex scenarios is solved, and higher safety and the implementation of autonomous driving technology are achieved.
Patent Information
- Application Number
- CN202311735181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-12-15
AI Technical Summary
The generalization and interaction of autonomous driving models in complex real-world scenarios lead to safety hazards and floor-based obstacles.
By constructing a visual Q&A network diagram, the driving behavior prediction model and driving trajectory prediction model are trained based on the autonomous driving data set and/or the simulator to improve the generalization and interactivity of the model.
It improves the generalization ability and interactivity of autonomous driving models in complex scenarios, reduces safety risks, and promotes the implementation of autonomous driving technology.
Smart Images

Figure CN117668761B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of autonomous driving, and in particular, to a method, device, equipment and storage medium for training an autonomous driving model. Background Art
[0002] Trajectory planning is a core technology of an autonomous driving system. At present, a large number of neural network models are used for trajectory planning in the field of autonomous driving. However, the lack of generalization has always been a major problem faced by the autonomous driving field. Models trained with relatively single and simple datasets cannot generalize well in complex real-world scenarios, which has led to many potential safety hazards and seriously hindered the implementation of autonomous driving. At the same time, an autonomous driving model also requires a certain degree of interactivity. Therefore, it is particularly important to make up for the deficiencies in generalization and interactivity in the field of autonomous driving. Summary of the Invention
[0003] The present invention provides a method, device, equipment and storage medium for training an autonomous driving model. By training the autonomous driving model based on the constructed visual question answering network graph, the generalization and interactivity of the autonomous driving model can be improved.
[0004] In a first aspect, an embodiment of the present invention provides a method for training an autonomous driving model. The autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model. The method includes:
[0005] Constructing a visual question answering network graph based on an autonomous driving dataset and / or an autonomous driving simulator; wherein, the visual question answering network graph includes image frames and their corresponding question answering network graphs. The question answering network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question and answer pair;
[0006] Using the visual question answering network graph as a sample set, and dividing the sample set into a training set and a test set;
[0007] Training the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set.
[0008] In a second aspect, an embodiment of the present invention further provides a device for training an autonomous driving model. The autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model. The device includes:
[0009] A visual question answering network graph construction module, configured to construct a visual question answering network graph based on an autonomous driving data set and / or an autonomous driving simulator; wherein, the visual question answering network graph includes an image frame and its corresponding question answering network graph, the question answering network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question and answer pair;
[0010] A sample set division module, configured to use the visual question answering network graph as a sample set, and divide the sample set into a training set and a test set;
[0011] A model training and evaluation module, configured to train the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluate the driving behavior prediction model and the driving trajectory prediction model based on the test set.
[0012] In a third aspect, an embodiment of the present invention further provides an electronic device, the electronic device includes:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the autonomous driving model according to the embodiment of the present invention.
[0016] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the training method of the autonomous driving model according to the embodiment of the present invention when executed by a processor.
[0017] An embodiment of the present invention discloses a training method, device, equipment and storage medium for an autonomous driving model. The autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model. The method includes: constructing a visual question answering network graph based on an autonomous driving data set and / or an autonomous driving simulator; wherein, the visual question answering network graph includes image frames and corresponding question answering network graphs, and the question answering network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question and answer pair; using the visual question answering network graph as a sample set, and dividing the sample set into a training set and a test set; training the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set. The training method for the autonomous driving model provided by the embodiment of the present invention can improve the generalization and interactivity of the autonomous driving model by training and evaluating the autonomous driving model based on the visual question answering network graph constructed by the autonomous driving data set and / or the autonomous driving simulator. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of a training method for an autonomous driving model in Embodiment 1 of the present invention;
[0019] Figure 2 is an example diagram of a question answering network graph in Embodiment 1 of the present invention;
[0020] Figure 3 is a schematic structural diagram of a training device for an autonomous driving model in Embodiment 2 of the present invention;
[0021] Figure 4 is a schematic structural diagram of an electronic device in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only parts related to the present invention are shown in the drawings for the sake of convenience of description, rather than all structures.
[0023] Embodiment 1
[0024] Figure 1 is a flowchart of a training method for an autonomous driving model provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of training an autonomous driving model. The method can be executed by a training device for an autonomous driving model. The device can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC or a server, etc.
[0025] In this embodiment, the autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model. Among them, the driving behavior prediction model is used to predict the driving behavior of the vehicle itself, and the driving behavior may include driving speed and / or steering angle. The driving trajectory prediction model is used to predict the driving trajectory of the vehicle itself within the next n seconds, and the driving trajectory can be represented by the position coordinates of multiple trajectory points.
[0026] As Figure 1 shown, the method specifically includes the following steps:
[0027] S110, construct a visual question answering network graph based on the autonomous driving dataset and / or the autonomous driving simulator.
[0028] Among them, the visual question answering network graph includes image frames and corresponding question answering network graphs. The question answering network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question and answer pair (QA). Each node (which can be called a QA node) corresponds to an inference stage. If there is a directed edge between the nodes, it indicates that there is a logical dependency between the inference stages corresponding to the two nodes.
[0029] Among them, the inference stage may include at least one of the perception stage, the prediction stage, and the planning stage. Perception can be understood as identifying, describing, and localizing objects in the current driving scene; prediction can be understood as estimating possible actions or interactions of objects based on the perception results; planning can be understood as possible safe actions of the vehicle itself.
[0030] In this embodiment, the directed edge can correspond to logical dependency relationships in two dimensions, namely the object level dimension and the task level dimension. On the object level dimension, the mutual influence between different objects can be represented by the directed edge. Exemplarily, Figure 2 is an example diagram of a question answering network graph in an invention embodiment. As Figure 2 shown, if there is a directed edge between the planning node (QA6) of the sedan 2 and the perception node (QA7) of the pedestrian, it indicates that the planning node of the sedan 2 will be affected by the perception node of the pedestrian. On the task level dimension, the logical progression relationship between different inference stages is represented by the directed edge, such as Figure 2 shown, progressing from the perception stage to the prediction stage, and then from the prediction stage to the planning stage.
[0031] Among them, the autonomous driving dataset can be an open-source dataset for the field of autonomous driving, such as the nuScenes dataset. The autonomous driving simulator can be an open-source tool for simulating autonomous driving, such as the CARLA simulator.
[0032] In this embodiment, the process of constructing a visual question-and-answer network graph based on an autonomous driving dataset can be as follows: Extract key frames from the driving videos in the autonomous driving dataset; Extract at least one key object from the key frames; Add question-and-answer pairs for multiple inference stages to each key object; Add directed edges between the question-and-answer pairs based on the logical dependency relationships between the question-and-answer pairs to obtain the question-and-answer network graph corresponding to the key frame, and add 2D bounding boxes to the key objects in the key frame.
[0033] Among them, the inference stages include at least one of the perception stage, the prediction stage, and the planning stage. For the detailed elaboration of each inference stage, please refer to the above embodiments and will not be elaborated here. The driving videos can be collected by a camera installed on the vehicle itself. The autonomous driving dataset contains a large number of driving videos, and several driving videos can be selected from them for the solution of this embodiment. The key object can be other driving objects in the key frame that may affect the driving decision of the vehicle itself, such as vehicles, pedestrians, etc. Specifically, the method of extracting key frames from the driving videos in the autonomous driving dataset can be: Input the driving video into a pre-trained key frame extraction neural network model and output the key frames contained in the driving video; Or, the annotator screens out the key frames from the driving video. The method of extracting at least one key object from the key frames can be: Input the key frame into a pre-trained key extraction neural network model; Or the annotator extracts the key object from the key frame.
[0034] Among them, there can be one or more question-and-answer pairs for each key object in each inference stage, that is, each key object contains question-and-answer pairs of the perception stage - prediction stage - planning stage. Specifically, the process of adding question-and-answer pairs for multiple inference stages to each key object can be: For the question-and-answer pairs in the perception stage, some of them can be generated according to the driving-related data in the autonomous driving dataset, and the other part can be manually annotated by the annotator; For the question-and-answer pairs in the prediction stage and the planning stage, they are manually annotated by the annotator.
[0035] Among them, the method of adding directed edges between the question-and-answer pairs based on the logical dependency relationships between the question-and-answer pairs can be: Establish the logical dependency relationships between the question-and-answer pairs according to the requirements, and add directed edges between the question-and-answer pairs according to the logical dependency relationships to obtain the question-and-answer network graph corresponding to the key frame. The method of adding 2D bounding boxes to the key objects in the key frame can be: Frame the key objects in the key frame with 2D bounding boxes and annotate the position information of the 2D bounding boxes. Among them, the position information of the 2D bounding boxes can be represented by the position coordinates of the two vertices on the diagonal. That is, the key frame with the added 2D bounding box is used as the image frame in the visual question-and-answer network graph. Optionally, in this embodiment, a scene-level description can also be added to each key frame, and this scene-level description elaborates on the behavior of the vehicle itself in the entire driving video.
[0036] In this embodiment, the method for constructing a visual question-and-answer network graph based on an autonomous driving simulator may be as follows: create a virtual ego vehicle and a virtual scene based on the autonomous driving simulator; control the virtual ego vehicle to drive in the virtual scene, and collect driving-related data of the virtual ego vehicle during driving; construct question-and-answer pairs for multiple inference stages corresponding to the image frames based on the driving-related data; add directed edges between the question-and-answer pairs based on the logical dependency relationships between the question-and-answer pairs to obtain the question-and-answer network graph corresponding to the image frames.
[0037] Among them, the driving-related data can be collected through sensors installed in the autonomous driving simulator, including: semantic segmentation, depth maps, lidar point cloud data, etc.
[0038] In this embodiment, the autonomous driving simulator can be used to collect data in Leaderboard 2.0, and a rule-based virtual ego vehicle with priorities can be adopted. Among them, Leaderboard 2.0 introduces two new large maps and a new set of scenarios, enhancing the diversity of the training and evaluation environments. One of the virtual scenarios (e.g., Town 12) can be used to construct a training set, and another virtual scenario (e.g., Town 13) can be used to construct a test set. A series of routes are set in the urban, residential, and rural areas of the virtual scenario, enabling the virtual ego vehicle to drive along these routes.
[0039] Specifically, the process of constructing question-and-answer pairs for multiple inference stages corresponding to the image frames based on the driving-related data can be as follows: based on the inquiry statements about road layout, stop signs, traffic lights, and vehicles included in the original data set in the autonomous driving simulator, combined with the collected driving-related data, annotators generate question-and-answer pairs for multiple inference stages corresponding to the image frames.
[0040] Among them, the method of adding directed edges between the question-and-answer pairs based on the logical dependency relationships between the question-and-answer pairs can be: establish the logical dependency relationships between the question-and-answer pairs according to requirements, and add directed edges between the question-and-answer pairs based on the logical dependency relationships to obtain the question-and-answer network graph corresponding to the key frames.
[0041] S120: Use the visual question-and-answer network graph as a sample set, and divide the sample set into a training set and a test set.
[0042] Among them, the method of dividing the sample set into a training set and a test set can be: divide the sample set into a training set and a test set according to a certain ratio. For example: 8:2.
[0043] S130: Train the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluate the driving behavior prediction model and the driving trajectory prediction model based on the test set.
[0044] Among them, the driving behavior prediction model and the driving trajectory prediction model can be constructed based on the Vision-Language Model (VLM). In this embodiment, the question-answer pairs at different inference stages in the visual question-answering network graph in the training set are used as the inputs of the driving behavior prediction model and the driving trajectory prediction model for model training.
[0045] Among them, the question-answer pair includes an interrogation statement and a true response statement. The visual question-answering network graph also carries the true driving behavior and the true driving trajectory corresponding to the image frame. The true driving behavior includes the true driving speed and the true steering angle. The driving trajectory is characterized by the position coordinates of multiple trajectory points.
[0046] In this embodiment, the method for training the driving behavior prediction model based on the training set can be: traversing each node in turn according to the logical dependency relationship between the nodes in the visual question-answering network graph until all nodes are traversed: if the traversed node has no previous node, the image frame and the interrogation statement of the node are input into the driving behavior prediction model, and the predicted response statement corresponding to the node is output; if the traversed node has a previous node, the question-answer pair of the previous node and the interrogation statement of the node are input into the driving behavior prediction model, and the predicted response statement of the node is output; the question-answer pairs of all nodes in the visual question-answering network graph are input into the driving behavior prediction model to output the predicted driving behavior; the driving behavior prediction model is trained based on the predicted response statements of each node, the true response statements of each node, the predicted driving behavior, and the true driving behavior.
[0047] Among them, the driving behavior can include two pieces of information: driving speed and steering angle. Among them, the driving speed can include: very slow, slow, normal speed, fast, and very fast, five speed levels, and the speed ranges corresponding to each speed level can be preset in advance. The steering angle can include five steering levels: left turn, slight left turn, straight ahead, slight right turn, and right turn, and the angle ranges corresponding to each steering level can be preset.
[0048] Among them, traversing each node in turn according to the logical dependency relationship between the nodes in the visual question-answering network graph can be understood as: starting from the node in the initial inference stage and traversing the nodes until the node in the ending inference stage is traversed. For example: in this embodiment, starting from the node in the perception stage and traversing the question-answering network graph until the node in the planning stage is traversed.
[0049] If the traversed node has no previous node, it indicates that the node is in the starting inference stage (such as the perception stage). Then, the image frame and the interrogation statement of the node are input into the driving behavior prediction model, and the predicted response statement of the node is output. If the traversed node has previous nodes, it indicates that the node is in the intermediate or ending inference stage (such as the prediction stage or the planning stage). Then, the Q&A pairs of all the previous nodes of the node and the interrogation statement of the node are input into the driving behavior prediction model, and the predicted response statement of the node is output. Exemplarily, as Figure 2 shown, QA1, QA4, and QA7 are nodes without previous nodes, and the remaining nodes are nodes with previous nodes. For QA6, its previous nodes include QA4, QA5, and QA7. The Q&A pairs of all the nodes in the visual Q&A network graph are input into the driving behavior prediction model, and the predicted driving behavior is output.
[0050] Specifically, the way to train the driving behavior prediction model based on the predicted response statements of each node, the true response statements of each node, the predicted driving behavior, and the true driving behavior can be as follows: Determine the first sub-loss function according to the predicted response statements of each node and the corresponding true response statements, determine the second sub-loss function according to the predicted driving behavior and the true driving behavior of the visual Q&A network graph, superimpose the first sub-loss function and the second sub-loss function to obtain the final loss function; finally, train the driving behavior prediction model based on the final loss function.
[0051] In this embodiment, the way to train the driving trajectory prediction model based on the training set can be: Input the image frame of the visual Q&A network graph, the Q&A pairs of all its nodes, and the true driving behavior into the driving trajectory prediction model, and output the predicted driving trajectory; train the driving trajectory prediction model based on the predicted driving trajectory and the true driving trajectory.
[0052] Among them, since the driving trajectory prediction model cannot output fine numerical results, in this embodiment, trajectory word segmentation can be used to represent the coordinate values. According to the statistical data of the training set trajectories, both the horizontal axis and the vertical axis of the driving trajectory can be divided into a set number (such as 256) of partitions, and then the word segmentation in the word segmenter is redefined to establish a corresponding relationship between each partition on the coordinate axis and the word segmentation, so that the driving trajectory prediction model can represent the driving trajectory by outputting word segmentation. Specifically, input the image frame of the visual Q&A network graph, the Q&A pairs of all its nodes, and the true driving behavior into the driving trajectory prediction model, and output the predicted driving trajectory.
[0053] Among them, the way to train the driving trajectory prediction model based on the predicted driving trajectory and the true driving trajectory can be: Determine the loss function according to the predicted driving trajectory and the true driving trajectory, and train the driving trajectory prediction model based on this loss function.
[0054] In this embodiment, after training the driving behavior prediction model and the driving trajectory prediction model based on the training set, the driving behavior prediction model and the driving trajectory prediction model are evaluated based on the test set to evaluate the accuracy of the driving behavior prediction model and the driving trajectory prediction model.
[0055] Specifically, the process of evaluating the driving behavior prediction model based on the test set can be as follows: Traverse each node in sequence according to the logical dependencies between the nodes in the visual question and answer network diagram until all nodes are traversed. If the traversed node has no previous node, input the image frame and the query statement of the node into the trained driving behavior prediction model, and output the predicted response statement corresponding to the node. If the traversed node has a previous node, input the Q&A pair of the previous node and the query statement of the node into the trained driving behavior prediction model, and output the predicted response statement of the node. Input the Q&A pairs of all nodes in the visual question and answer network diagram into the driving behavior prediction model, and output the predicted driving behavior. Evaluate the predicted response statement based on the true response statement of each node to obtain the Q&A evaluation index, and evaluate the predicted driving behavior based on the true driving behavior to obtain the accuracy evaluation index.
[0056] Among them, the accuracy evaluation index includes the speed accuracy evaluation index and the steering angle accuracy evaluation index. The Q&A evaluation index can include the SPICE (Semantic Propositional Image Caption Evaluation) index and / or the GPT (Generative Pre-trained Transformer) index. In this embodiment, after the driving behavior prediction model is trained, the Q&A pairs of the visual question and answer network diagram in the test set are input into the driving behavior prediction model in the above manner, and the predicted driving behavior corresponding to each node and the predicted driving behavior corresponding to each visual question and answer network diagram are output. Evaluate the predicted response statement based on the true response statement of each node to obtain the SPICE index and the GPT index of the trained driving behavior prediction model. Evaluate the predicted driving behavior based on the true driving behavior to obtain the speed accuracy and the steering angle accuracy of the trained driving behavior prediction model.
[0057] Specifically, the method of evaluating the driving trajectory prediction model based on the test set can be: Input the Q&A pairs of all nodes in the visual question and answer network diagram and the true driving behavior into the trained driving trajectory prediction model, and output the predicted driving trajectory. Evaluate the predicted driving trajectory based on the true driving trajectory to obtain the error evaluation index.
[0058] Among them, the error evaluation metrics include the average displacement error evaluation metric and the final displacement error evaluation metric. The average displacement error evaluation metric can be characterized by the average Euclidean distance between the predicted driving trajectory and the true driving trajectory; the final error evaluation metric can be characterized by the Euclidean distance between the predicted end point and the true end point.
[0059] Optionally, after evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set, the following steps are further included: obtaining a target image frame and a plurality of interrogation sentences; sequentially inputting the target image and the plurality of interrogation sentences into the driving behavior prediction model according to the logical relationship to obtain the target driving behavior and the response sentences of each interrogation sentence; inputting the target image frame, the plurality of question-and-answer pairs, and the target driving behavior into the driving trajectory prediction model, and outputting the target driving trajectory corresponding to the target image frame.
[0060] Among them, there is a logical relationship among the plurality of interrogation sentences; the interrogation sentence and the response sentence form a question-and-answer pair. The target image frame can be an image of the surrounding environment collected during the driving of the vehicle. The plurality of interrogation sentences can be logically progressive in different inference stages, that is, they have a logical relationship. Specifically, first input the target image frame and the first interrogation sentence into the driving behavior prediction model; output the first response sentence; then input the first question-and-answer pair and the second interrogation sentence into the driving behavior prediction model, and output the second response sentence; then input the first question-and-answer pair, the second question-and-answer pair, and the third interrogation sentence into the driving behavior prediction model, and output the third response sentence; and so on, until all the interrogation sentences are input into the driving behavior prediction model to obtain the target driving behavior and a plurality of question-and-answer pairs. Finally, input the target image frame, the plurality of question-and-answer pairs, and the target driving behavior into the driving trajectory prediction model, and output the target driving trajectory corresponding to the target image frame.
[0061] The technical solution of this embodiment constructs a visual question-and-answer network graph based on an autonomous driving data set and / or an autonomous driving simulator; among them, the visual question-and-answer network graph includes an image frame and a corresponding question-and-answer network graph, the question-and-answer network graph is composed of a plurality of nodes and directed edges between the nodes, and each node carries a question-and-answer pair; using the visual question-and-answer network graph as a sample set, and dividing the sample set into a training set and a test set; training the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set. The training method of the autonomous driving model provided by the embodiment of the present invention can improve the generalization and interactivity of the autonomous driving model by training and evaluating the autonomous driving model based on the visual question-and-answer network graph constructed by the autonomous driving data set and / or the autonomous driving simulator.
[0062] Embodiment 2
[0063] Figure 3It is a schematic structural diagram of a training device for an autonomous driving model provided in the second embodiment of the present invention. The autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model. The device includes:
[0064] A visual question-and-answer network graph construction module 310, configured to construct a visual question-and-answer network graph based on an autonomous driving data set and / or an autonomous driving simulator; wherein, the visual question-and-answer network graph includes an image frame and its corresponding question-and-answer network graph. The question-and-answer network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question-and-answer pair.
[0065] A sample set division module 320, configured to use the visual question-and-answer network graph as a sample set and divide the sample set into a training set and a test set.
[0066] A model training and evaluation module 330, configured to train the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluate the driving behavior prediction model and the driving trajectory prediction model based on the test set.
[0067] Optionally, the visual question-and-answer network graph construction module 310 is further configured to:
[0068] Extract key frames from the driving video of the autonomous driving data set;
[0069] Extract at least one key object from the key frames;
[0070] Add question-and-answer pairs for multiple inference stages to each key object; wherein, the inference stages include at least one of a perception stage, a prediction stage, and a planning stage;
[0071] Add directed edges between the question-and-answer pairs based on the logical dependency relationship between the question-and-answer pairs to obtain the question-and-answer network graph corresponding to the key frames, and add 2D bounding boxes to the key objects in the key frames.
[0072] Optionally, the visual question-and-answer network graph construction module 310 is further configured to:
[0073] Create a virtual ego vehicle and a virtual scene based on the autonomous driving simulator;
[0074] Control the virtual ego vehicle to drive in the virtual scene and collect driving-related data during the driving process of the virtual ego vehicle;
[0075] Construct question-and-answer pairs for multiple inference stages corresponding to the image frames based on the driving-related data;
[0076] Add directed edges between the question-and-answer pairs based on the logical dependency relationship between the question-and-answer pairs to obtain the question-and-answer network graph corresponding to the image frames.
[0077] Optionally, the question-and-answer pair includes an interrogation statement and a true response statement; the visual question-and-answer network diagram also carries the true driving behavior and true driving trajectory corresponding to the image frame; wherein, the true driving behavior includes the true driving speed and the true steering angle.
[0078] Optionally, the model training and evaluation module 330 is further configured to:
[0079] Traverse each node in sequence according to the logical dependency relationship between the nodes in the visual question-and-answer network diagram until all nodes are traversed: if the traversed node has no previous node, input the image frame and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement corresponding to the node; if the traversed node has a previous node, input the question-and-answer pair of the previous node and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement of the node;
[0080] Input the question-and-answer pairs of all nodes in the visual question-and-answer network diagram into the driving behavior prediction model, and output the predicted driving behavior;
[0081] Train the driving behavior prediction model based on the predicted response statements of each node, the true response statements of each node, the predicted driving behavior, and the true driving behavior.
[0082] Optionally, the model training and evaluation module 330 is further configured to:
[0083] Input the image frame of the visual question-and-answer network diagram, the question-and-answer pairs of all its nodes, and the true driving behavior into the driving trajectory prediction model, and output the predicted driving trajectory;
[0084] Train the driving trajectory prediction model based on the predicted driving trajectory and the true driving trajectory.
[0085] Optionally, the model training and evaluation module 330 is further configured to:
[0086] Traverse each node in sequence according to the logical dependency relationship between the nodes in the visual question-and-answer network diagram until all nodes are traversed: if the traversed node has no previous node, input the image frame and the interrogation statement of the node into the trained driving behavior prediction model, and output the predicted response statement corresponding to the node; if the traversed node has a previous node, input the question-and-answer pair of the previous node and the interrogation statement of the node into the trained driving behavior prediction model, and output the predicted response statement of the node;
[0087] Input the question-and-answer pairs of all nodes in the visual question-and-answer network diagram into the driving behavior prediction model, and output the predicted driving behavior;
[0088] Evaluate the predicted response statement based on the true response statements of each node to obtain a Q&A evaluation metric, and evaluate the predicted driving behavior based on the true driving behavior to obtain an accuracy evaluation metric; wherein, the accuracy evaluation metric includes a speed accuracy evaluation metric and a steering angle accuracy evaluation metric.
[0089] Optionally, the model training and evaluation module 330 is further configured to:
[0090] Input the Q&A pairs of all nodes in the visual Q&A network graph and the true driving behavior into the trained driving trajectory prediction model, and output the predicted driving trajectory;
[0091] Evaluate the predicted driving trajectory based on the true driving trajectory to obtain an error evaluation metric; wherein, the error evaluation metric includes an average displacement error evaluation metric and a final displacement error evaluation metric.
[0092] Optionally, it further includes: a target driving trajectory determination module, configured to:
[0093] Obtain a target image frame and a plurality of query statements; wherein, there is a logical relationship among the plurality of query statements;
[0094] Input the target image and the plurality of query statements into the driving behavior prediction model in sequence according to the logical relationship to obtain the target driving behavior and the response statements of each query statement; wherein, the query statement and the response statement form a Q&A pair;
[0095] Input the target image frame, the plurality of Q&A pairs and the target driving behavior into the driving trajectory prediction model, and output the target driving trajectory corresponding to the target image frame.
[0096] The above device can execute the methods provided in all the foregoing embodiments of the present invention, and has corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of the present invention.
[0097] Embodiment III
[0098] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0099] As shown Figure 4 in FIG. 1, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0100] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0101] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the training method of an autonomous driving model.
[0102] In some embodiments, the training method of the autonomous driving model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method of the autonomous driving model described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the training method of the autonomous driving model by any other appropriate means (e.g., by means of firmware).
[0103] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0104] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0105] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0107] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0108] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0109] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0110] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for training an autonomous driving model, characterized in that, the autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model, and the method includes: constructing a visual question answering network graph based on an autonomous driving data set and / or an autonomous driving simulator; wherein, the visual question answering network graph includes image frames and their corresponding question answering network graphs, the question answering network graph is composed of multiple nodes and directed edges between the nodes, and each node carries a question and answer pair; using the visual question answering network graph as a sample set, and dividing the sample set into a training set and a test set; training the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set; the question and answer pair includes an interrogation statement and a true response statement; the visual question answering network graph also carries the true driving behavior and true driving trajectory corresponding to the image frame; wherein, the true driving behavior includes a true driving speed and a true steering angle; training the driving behavior prediction model based on the training set includes: traversing each node in sequence according to the logical dependency relationship between the nodes in the visual question answering network graph until all nodes are traversed: if the traversed node has no previous node, input the image frame and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement corresponding to the node; if the traversed node has a previous node, input the question and answer pair of the previous node and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement of the node; inputting the question and answer pairs of all nodes in the visual question answering network graph into the driving behavior prediction model, and outputting a predicted driving behavior; training the driving behavior prediction model based on the predicted response statements of the respective nodes, the true response statements of the respective nodes, the predicted driving behavior, and the true driving behavior.
2. The method according to claim 1, characterized in that, constructing a visual question answering network graph based on an autonomous driving data set includes: extracting key frames from the driving video of the autonomous driving data set; extracting at least one key object from the key frames; adding question and answer pairs for multiple inference stages to each key object; wherein, the inference stages include at least one of a perception stage, a prediction stage, and a planning stage; adding directed edges between the question and answer pairs based on the logical dependency relationship between the question and answer pairs, obtaining the question answering network graph corresponding to the key frame, and adding 2D bounding boxes to the key objects in the key frame.
3. The method according to claim 1, characterized in that, constructing a visual question answering network graph based on an autonomous driving simulator includes: creating a virtual ego vehicle and a virtual scene based on the autonomous driving simulator; controlling the virtual ego vehicle to drive in the virtual scene, and collecting driving-related data of the virtual ego vehicle during driving; constructing question and answer pairs for multiple inference stages corresponding to the image frame based on the driving-related data; Adding directed edges between the question-and-answer pairs based on the logical dependency relationships between the question-and-answer pairs to obtain a question-and-answer network graph corresponding to the image frame.
4. The method according to claim 1, wherein, training the driving trajectory prediction model based on the training set includes: inputting the image frame of the visual question-and-answer network graph, the question-and-answer pairs of all its nodes, and the real driving behavior into the driving trajectory prediction model, and outputting a predicted driving trajectory; training the driving trajectory prediction model based on the predicted driving trajectory and the real driving trajectory.
5. The method according to claim 1, wherein, evaluating the driving behavior prediction model based on the test set includes: traversing each node in sequence according to the logical dependency relationships between the nodes in the visual question-and-answer network graph until all nodes are traversed: if the traversed node has no previous nodes, inputting the image frame and the query statement of the node into the trained driving behavior prediction model, and outputting a predicted response statement corresponding to the node; if the traversed node has previous nodes, inputting the question-and-answer pair of the previous node and the query statement of the node into the trained driving behavior prediction model, and outputting a predicted response statement of the node; inputting the question-and-answer pairs of all nodes in the visual question-and-answer network graph into the driving behavior prediction model, and outputting a predicted driving behavior; evaluating the predicted response statements based on the real response statements of each node to obtain a question-and-answer evaluation index, and evaluating the predicted driving behavior based on the real driving behavior to obtain an accuracy evaluation index; wherein, the accuracy evaluation index includes a speed accuracy evaluation index and a steering angle accuracy evaluation index.
6. The method according to claim 1, wherein, evaluating the driving trajectory prediction model based on the test set includes: inputting the question-and-answer pairs of all nodes in the visual question-and-answer network graph and the real driving behavior into the trained driving trajectory prediction model, and outputting a predicted driving trajectory; evaluating the predicted driving trajectory based on the real driving trajectory to obtain an error evaluation index; wherein, the error evaluation index includes an average displacement error evaluation index and a final displacement error evaluation index.
7. The method according to claim 1, wherein, after evaluating the driving behavior prediction model and the driving trajectory prediction model based on the test set, further includes: acquiring a target image frame and a plurality of query statements; wherein, there is a logical relationship between the plurality of query statements; inputting the target image and the plurality of query statements into the driving behavior prediction model in sequence according to the logical relationship to obtain a target driving behavior and response statements of each query statement; wherein, the query statement and the response statement form a question-and-answer pair; inputting the target image frame, a plurality of the question-and-answer pairs, and the target driving behavior into the driving trajectory prediction model, and outputting a target driving trajectory corresponding to the target image frame.
8. A training device for an autonomous driving model, wherein, The autonomous driving model includes a driving behavior prediction model and a driving trajectory prediction model, and the device includes: a visual question answering network graph construction module, configured to construct a visual question answering network graph based on an autonomous driving data set and / or an autonomous driving simulator; wherein, the visual question answering network graph includes an image frame and its corresponding question answering network graph, the question answering network graph is composed of a plurality of nodes and directed edges between the nodes, and each node carries a question and answer pair; a sample set division module, configured to use the visual question answering network graph as a sample set, and divide the sample set into a training set and a test set; a model training and evaluation module, configured to train the driving behavior prediction model and the driving trajectory prediction model based on the training set, and evaluate the driving behavior prediction model and the driving trajectory prediction model based on the test set; The question and answer pair includes an interrogation statement and a true response statement; the visual question answering network graph also carries the true driving behavior and the true driving trajectory corresponding to the image frame; wherein, the true driving behavior includes a true driving speed and a true steering angle; The model training and evaluation module is further configured to: traverse each node in sequence according to the logical dependency relationship between the nodes in the visual question answering network graph until all nodes are traversed: if the traversed node has no previous node, input the image frame and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement corresponding to the node; if the traversed node has a previous node, input the question and answer pair of the previous node and the interrogation statement of the node into the driving behavior prediction model, and output the predicted response statement of the node; input the question and answer pairs of all nodes in the visual question answering network graph into the driving behavior prediction model, and output the predicted driving behavior; train the driving behavior prediction model based on the predicted response statements of each node, the true response statements of each node, the predicted driving behavior and the true driving behavior.
9. An electronic device, characterized in that the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the training method of the autonomous driving model according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the training method of the autonomous driving model according to any one of claims 1-7 is realized.