Scene characterization model training method, obstacle marking method and autonomous vehicle
By training a scene representation model in autonomous vehicles and utilizing scene information from different driving scenarios, especially the differences in the distribution of key obstacles, the accuracy of scene representation and obstacle labeling is improved, solving the problem of insufficient accuracy in existing technologies and enhancing the behavioral decision-making capabilities of autonomous driving.
Patent Information
- Application Number
- CN202511375677.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2026-01-13
AI Technical Summary
In existing autonomous driving technologies, the computational logic set by humans results in low accuracy of scene representation and obstacle marking results, making them unsuitable for flexible application in various real-world driving scenarios.
By acquiring scene information from different driving scenarios, especially scene information containing differences in the distribution of key obstacles, the scene representation model is trained, and the accuracy of obstacle marking is improved by using a learnable neural network model.
It improves the accuracy of scene representation and obstacle labeling results, enhances the perception and processing capabilities of autonomous vehicles for key obstacles, and improves the accuracy of behavioral decisions.
Smart Images

Figure CN121328642A_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese invention entitled “Scene Representation Model Training Method, Obstacle Marking Method and Autonomous Vehicle”, application number “202310430715.X”, filed on April 19, 2023. Technical Field
[0002] This disclosure relates to the field of artificial intelligence, and more particularly to the field of autonomous driving technology, specifically to a scene representation model training method, an obstacle marking method, and an autonomous vehicle. Background Technology
[0003] Autonomous driving technology involves multiple aspects, including environmental perception, behavioral decision-making, trajectory planning, and motion control. Among these, behavioral decision-making relies on the scene representation results and obstacle marking results of the driving scene. Currently, the scene representation results and obstacle marking results are mainly obtained by using manually set computational logic.
[0004] Because the computational logic set by humans has significant limitations and cannot be flexibly applied to various actual driving scenarios, it will affect the accuracy of scene representation results and obstacle marking results. Summary of the Invention
[0005] This disclosure provides a method for training a scene representation model, an obstacle marking method, and an autonomous vehicle.
[0006] According to one aspect of this disclosure, a method for training a scene representation model is provided, comprising:
[0007] Obtain the first scene information of the first driving scenario;
[0008] Obtain second scenario information for the second driving scenario; wherein, the distribution of key obstacles differs between the second driving scenario and the first driving scenario, and the key obstacles are used to change the driving risk status of the main vehicle;
[0009] The scene representation model is trained based on the first scene information and the second scene information.
[0010] According to another aspect of this disclosure, an obstacle marking method is provided, comprising:
[0011] Obtain driving scene information for the current driving scenario;
[0012] Driving scene information is input into a trained scene representation model to obtain intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation based on the driving scene information; wherein, the scene representation model is trained by any scene representation model training method.
[0013] Based on intermediate parameters, the current obstacles in the current driving scene are marked to obtain the obstacle marking results.
[0014] According to another aspect of this disclosure, a scene representation model training apparatus is provided, comprising:
[0015] The first information acquisition unit is used to acquire the first scene information of the first driving scenario;
[0016] The second information acquisition unit is used to acquire second scene information of the second driving scenario; wherein, the second driving scenario differs from the first driving scenario in the distribution of key obstacles, and the key obstacles are used to change the driving risk state of the main vehicle;
[0017] The first model training unit is used to train the scene representation model based on the first scene information and the second scene information.
[0018] According to another aspect of this disclosure, an obstacle marking device is provided, comprising:
[0019] The current information acquisition unit is used to acquire driving scene information of the current driving scenario;
[0020] The intermediate parameter acquisition unit is used to input driving scene information into the trained scene representation model in order to obtain the intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation based on the driving scene information; wherein, the scene representation model is obtained by any scene representation model training method.
[0021] The obstacle marking unit is used to mark the current obstacles in the current driving scene based on intermediate parameters and obtain the obstacle marking results.
[0022] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0023] At least one processor;
[0024] The memory that is communicatively connected to the at least one processor;
[0025] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods in the embodiments of this disclosure.
[0026] According to another aspect of this disclosure, an autonomous vehicle is provided, including electronic devices.
[0027] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods in the embodiments of this disclosure.
[0028] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods in the embodiments of this disclosure.
[0029] Using this disclosure can improve the accuracy of scene characterization results and obstacle marking results.
[0030] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0031] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0032] Figure 1 A flowchart illustrating a scene representation model training method provided in an embodiment of this disclosure;
[0033] Figure 2A and 2B This diagram illustrates a derivative method of a second driving scenario provided by an embodiment of this disclosure.
[0034] Figure 3A and 3B This diagram illustrates another derivative method of a second driving scenario provided by an embodiment of this disclosure;
[0035] Figure 4A This is a schematic diagram of the network structure of a scene representation model provided in an embodiment of the present disclosure.
[0036] Figure 4B A schematic diagram of the network structure of another scene representation model provided in this embodiment of the disclosure;
[0037] Figure 5A and 5B This diagram illustrates a derivative method of a fourth driving scenario provided by an embodiment of this disclosure.
[0038] Figure 6 A scene illustration of a scene representation model training method provided in this embodiment of the disclosure;
[0039] Figure 7 A flowchart illustrating an obstacle marking method provided in this embodiment of the present disclosure;
[0040] Figure 8 A schematic diagram of a scenario for an obstacle marking method provided in an embodiment of this disclosure;
[0041] Figure 9A schematic structural block diagram of a scene representation model training device provided in this disclosure embodiment;
[0042] Figure 10 A schematic structural block diagram of an obstacle marking device provided in an embodiment of this disclosure;
[0043] Figure 11 This is a schematic structural block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0044] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0045] This disclosure provides a method for training a scene representation model, which can be applied to electronic devices. The following will be combined with... Figure 1 The flowchart shown illustrates a method for training a scene representation model according to an embodiment of this disclosure. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order.
[0046] Step S101: Obtain the first scene information of the first driving scenario;
[0047] Step S102: Obtain second scenario information for the second driving scenario; wherein, the second driving scenario differs from the first driving scenario in the distribution of key obstacles, and the key obstacles are used to change the driving risk state of the main vehicle;
[0048] Step S103: Train the scene representation model based on the first scene information and the second scene information.
[0049] The first driving scenario can be a historical driving scenario of the main vehicle, and the second driving scenario can be a derived driving scenario obtained by adjusting the key obstacles in the first driving scenario. Therefore, the distribution of key obstacles differs between the second and first driving scenarios, and these key obstacles are used to change the driving risk state of the main vehicle. The driving risk state can be any one of a safe state (no collision risk and no emergency braking risk), a collision risk, and an emergency braking risk. Furthermore, it should be noted that in this embodiment, when the main vehicle is driving in a certain driving scenario, if the distance between the obstacle and the main vehicle is less than the safe distance threshold throughout the entire journey, the main vehicle is considered to have no collision risk; otherwise, the main vehicle is considered to have a collision risk. Similarly, when the main vehicle is driving in a certain driving scenario, if the main vehicle does not brake suddenly throughout the entire journey, the main vehicle is considered to have no emergency braking risk; otherwise, the main vehicle is considered to have an emergency braking risk. The safe distance threshold can be set according to actual application needs; for example, it can be set to 20 centimeters (cm), and this embodiment does not limit this setting.
[0050] In this embodiment of the disclosure, the first scene information may include first master vehicle information and first obstacle information. The first master vehicle information may include the master vehicle's speed, acceleration, and pose information in the first driving scene. The first obstacle information may include the obstacle type in the first driving scene, specifically a movable obstacle (e.g., an obstacle vehicle) or a fixed obstacle (e.g., a bridge pier). If the obstacle type is a movable obstacle, the first obstacle information may also include the obstacle's speed, acceleration, and pose information in the first driving scene. If the obstacle type is a fixed obstacle, the first obstacle information may also include the location of the area corresponding to the obstacle in the first driving scene.
[0051] Similarly, in this embodiment, the second scene information may include second master vehicle information and second obstacle information. The second master vehicle information may include the master vehicle's speed, acceleration, and pose information in the second driving scene. The second obstacle information may include the obstacle type in the second driving scene, specifically a movable obstacle or a fixed obstacle. If the obstacle type is a movable obstacle, the second obstacle information may also include the obstacle's speed, acceleration, and pose information in the second driving scene. If the obstacle type is a fixed obstacle, the second obstacle information may also include the location of the obstacle in the area corresponding to the obstacle in the second driving scene.
[0052] After obtaining the first scene information of the first driving scenario and the second scene information of the second driving scenario, the scene representation model can be trained based on the first scene information and the second scene information.
[0053] In a specific example, first scene information can be input into a scene representation model to obtain a first representation result output by the scene representation model; second scene information can be input into the scene representation model to obtain a second representation result output by the scene representation model; and the scene representation model can be trained based on the first and second representation results. The "training the scene representation model based on the first and second representation results" can include: obtaining a first loss between the first and second representation results; obtaining a trained scene representation model if the first loss meets the first loss requirement; and adjusting the parameters of the scene representation model if the first loss does not meet the first loss requirement. The first loss can be a first cosine similarity, and the first loss requirement can be that the first cosine similarity is less than a first similarity threshold. The first similarity threshold can be set according to actual application needs; for example, it can be set to 0.3, and this embodiment of the present disclosure does not limit this.
[0054] The scene representation model is a learnable neural network model. A trained scene representation model can be used to output the current scene representation result based on the input driving scene information. The driving scene information can be obtained during the actual driving process of the current vehicle, based on the current driving scene, and can have the same data structure as the first and second scene information mentioned above; details are omitted here.
[0055] The scene representation model training method provided in this disclosure can acquire first scene information of a first driving scenario; acquire second scene information of a second driving scenario; wherein the second driving scenario differs from the first driving scenario in the distribution of key obstacles, and these key obstacles are used to change the driving risk state of the main vehicle; and train the scene representation model based on the first and second scene information. After obtaining the trained scene representation model, it can be used to: output the current scene representation result based on the input driving scenario information. On the one hand, since the scene representation model is a learnable neural network model, compared with the prior art, obtaining the scene representation result based on the trained scene representation model can improve the accuracy of the scene representation result. On the other hand, since the distribution of key obstacles differs between the second driving scenario and the first driving scenario, and key obstacles are used to change the driving risk state of the main vehicle, key obstacles play a crucial role in the main vehicle's behavioral decisions and are important information for any driving scenario. Based on this, in this embodiment of the disclosure, by training the scene representation model with dissimilar scene pairs (first driving scenario and second driving scenario) that have different distributions of key obstacles, the scene representation model's ability to perceive and process key obstacles can be improved, thereby increasing the proportion and accuracy of key obstacle information in the scene representation results, and further improving the accuracy of the scene representation results.
[0056] As mentioned above, in this embodiment of the present disclosure, the first driving scenario can be a historical driving scenario of the main vehicle, and the second driving scenario can be a derived driving scenario obtained by adjusting the key obstacles in the first driving scenario. Based on this, in some optional embodiments, the scene representation model training method may further include the following steps:
[0057] Select the target obstacle from the first driving scenario;
[0058] Adjust the target obstacle in the first driving scenario to obtain the undetermined driving scenario;
[0059] Perform autonomous driving simulation on the master vehicle in the undetermined driving scenario to obtain the simulation risk state corresponding to the master vehicle when driving in the undetermined driving scenario;
[0060] When the simulated risk state differs from the historical risk state, the target obstacle is identified as the critical obstacle, and the pending driving scenario is identified as the second driving scenario; the historical risk state is the actual risk state corresponding to the main vehicle driving in the first driving scenario.
[0061] In this embodiment of the disclosure, a target obstacle can be selected from the first driving scenario based on the scenario category of the first driving scenario, or any obstacle can be selected from the first driving scenario as the target obstacle. The scenario category can include a safe scenario, a collision scenario, and an emergency braking scenario. A safe scenario is a driving scenario with no collision risk and no emergency braking risk, in which the driving risk state of the main vehicle is safe. A collision scenario is a driving scenario with collision risk, in which the driving risk state of the main vehicle is collision risk. An emergency braking scenario is a driving scenario with emergency braking risk, in which the driving risk state of the main vehicle is emergency braking risk.
[0062] The phrase "adjusting the target obstacle in the first driving scenario to obtain a pending driving scenario" can include: adjusting the pose and / or driving time of the target obstacle in the first driving scenario to obtain a pending driving scenario.
[0063] After obtaining the pending driving scenario, autonomous driving simulation software can be used to simulate autonomous driving of the main vehicle within the scenario, obtaining the simulated risk state corresponding to the main vehicle's operation in the pending driving scenario. If the simulated risk state differs from the historical risk state, the target obstacle is identified as a critical obstacle, and the pending driving scenario is designated as a secondary driving scenario. For example, if the simulated risk state is safe, but the historical risk state indicates a collision risk, the target obstacle is identified as a critical obstacle, and the pending driving scenario is designated as a secondary driving scenario.
[0064] Through the above steps, in this embodiment of the disclosure, a target obstacle can be selected from a first driving scenario, and the target obstacle can be adjusted in the first driving scenario to obtain a pending driving scenario. Then, autonomous driving simulation is performed on the main vehicle in the pending driving scenario to obtain the simulated risk state corresponding to the main vehicle driving in the pending driving scenario. This allows for the identification of the target obstacle as a critical obstacle and the determination of the pending driving scenario as a second driving scenario when the simulated risk state differs from the historical risk state. Since this embodiment of the disclosure obtains the simulated risk state corresponding to the main vehicle driving in the pending driving scenario through autonomous driving simulation, the efficiency of obtaining the simulated risk state can be improved, thereby improving the efficiency of identifying critical obstacles and thus improving the training efficiency of the scenario representation model.
[0065] In some alternative implementations, "selecting a target obstacle from the first driving scenario" may include the following steps:
[0066] Determine the scenario category for the first driving scenario;
[0067] In the case of a safe scenario, the first pending obstacle with the first posterior decision to be avoided is selected from the first driving scenario and is taken as the target obstacle; wherein, the first posterior decision is the driving decision made by the main vehicle to the first pending obstacle.
[0068] In the case of a collision scenario or an emergency braking scenario, any obstacle in the first driving scenario whose distance from the main vehicle is less than a first distance threshold is selected as the target obstacle.
[0069] That is, in this embodiment of the present disclosure, a target obstacle can be selected from the first driving scenario according to the scenario category of the first driving scenario.
[0070] In the case of a safety scenario, a first pending obstacle with the first posterior decision to be avoided is selected from the first driving scenario as the target obstacle. The avoidance can be lateral or longitudinal. Lateral avoidance can be left-turn or right-turn avoidance, and longitudinal avoidance can be acceleration avoidance in front or deceleration avoidance behind. This embodiment does not impose any limitations on these aspects.
[0071] In scenarios classified as collision scenarios or emergency braking scenarios, any obstacle in the first driving scenario whose distance from the main vehicle is less than a first distance threshold is selected as the target obstacle. The first distance threshold can be set according to actual application requirements; for example, it can be set to 30 meters (m). This embodiment of the present disclosure does not limit this setting.
[0072] Through the above steps, in this embodiment of the disclosure, a corresponding target obstacle determination strategy is provided for each type of first driving scenario, avoiding the use of a uniform target obstacle determination strategy. This ensures that the derived second driving scenario has a different scenario category from the first driving scenario, that is, it ensures that the main vehicle has different driving risk states in the first and second driving scenarios, thereby improving the training effect of the scenario representation model and further improving the accuracy of the scenario representation results.
[0073] In some alternative implementations, "adjusting the target obstacle in the first driving scenario to obtain the pending driving scenario" may include the following steps:
[0074] In the case of a safe scenario, the pose and / or travel time of the target obstacle are adjusted in the first driving scenario to obtain a pending driving scenario.
[0075] And / or, in the case of a collision scenario or an emergency braking scenario, remove the target obstacle in the first driving scenario to obtain a pending driving scenario.
[0076] Specifically, adjusting the pose of the target obstacle can involve adjusting its position and / or attitude to bring it closer to the main vehicle. Adjusting the travel time of the target obstacle, if it is a movable obstacle, can involve adjusting the time it enters the first driving scenario to bring it closer to the main vehicle. For example, if the target obstacle is in front of the main vehicle in the first driving scenario, its entry into the scenario can be controlled to be earlier; if it is behind the main vehicle, its entry time can be delayed.
[0077] Please combine Figure 2A and Figure 2B In the first driving scenario 201, the main vehicle 202 is included. If the scenario category of the first driving scenario 201 is a safe scenario, a first pending obstacle with the first posterior decision to be avoided can be selected from the first driving scenario 201 as the target obstacle 203. Subsequently, the pose and / or travel time of the target obstacle 203 can be adjusted in the first driving scenario 201 to obtain the pending driving scenario 204.
[0078] Please combine Figure 3A and Figure 3BIn the first driving scenario 301, a main vehicle 302 is included. If the scenario category of the first driving scenario 301 is a collision scenario or an emergency braking scenario, any obstacle in the first driving scenario 301 whose distance from the main vehicle is less than a first distance threshold can be selected as a target obstacle 303. Since the collision risk or emergency braking risk may be caused by the target obstacle 303, the target obstacle 303 can be directly removed in the first driving scenario 301 to obtain the pending driving scenario 304.
[0079] Through the above steps, in this embodiment of the disclosure, when the scenario category is a safety scenario, the pose and / or driving time of the target obstacle can be adjusted in the first driving scenario to obtain a pending driving scenario; when the scenario category is a collision scenario or an emergency braking scenario, the target obstacle can be directly removed in the first driving scenario to obtain a pending driving scenario, thereby simplifying the creation process of the pending driving scenario and further improving the training efficiency of the scenario representation model.
[0080] In some alternative implementations, "training the scene representation model based on the first scene information and the second scene information" may include the following steps:
[0081] The target scene information is input into the scene representation model to extract the fusion features of the main vehicle and the obstacle features from the target scene information; wherein, the target scene information is either the first scene information or the second scene information.
[0082] The scene representation model is obtained based on the correlation between the fusion features of the main vehicle and the obstacle features. Wherein, when the target scene information is the first scene information, the scene representation result is the first representation result, and when the target scene information is the second scene information, the scene representation result is the second representation result.
[0083] The scene representation model is trained based on the first and second representation results.
[0084] As described above, in this embodiment of the disclosure, when the target scene information is first scene information, the first scene information may include first main vehicle information and first obstacle information. In a specific example, multiple first sampling time points can be preset. The first main vehicle information may include the speed information, acceleration information, and pose information of the main vehicle collected from the first driving scene at each first sampling time point. The first obstacle information may include the obstacle type of any obstacle in the first driving scene, specifically a movable obstacle or a fixed obstacle. For each obstacle in the first driving scene, if the obstacle type is a movable obstacle, the first obstacle information may also include the speed information, acceleration information, and pose information of the obstacle collected from the first driving scene at each first sampling time point. If the obstacle type is a fixed obstacle, the first obstacle information may also include the area location corresponding to the obstacle in the first driving scene. The total number of first sampling time points can be set according to actual application requirements. For example, it can be set to 16, and two adjacent first sampling time points can be spaced 0.1 seconds (s). This embodiment of the disclosure does not limit this.
[0085] Similarly, in this embodiment of the disclosure, when the target scene information is second scene information, the second scene information may include second main vehicle information and second obstacle information. In a specific example, multiple second sampling time points can be preset. The second main vehicle information may include the speed information, acceleration information, and pose information of the main vehicle collected from the second driving scene at each second sampling time point. The second obstacle information may include the obstacle type of any obstacle in the second driving scene, specifically a movable obstacle or a fixed obstacle. For each obstacle in the second driving scene, if the obstacle type is a movable obstacle, the second obstacle information may also include the speed information, acceleration information, and pose information of the obstacle collected from the second driving scene at each second sampling time point. If the obstacle type is a fixed obstacle, the second obstacle information may also include the area location corresponding to the obstacle in the second driving scene. The total number of second sampling time points can be set according to actual application requirements. For example, it can be set to 16, and two adjacent second sampling time points can be spaced 0.1 seconds (s). This embodiment of the disclosure does not limit this.
[0086] Please combine Figure 4A For the scene representation model 400, in a specific example, it may include an encoder module 401 and a first attention module 402.
[0087] The encoder module 401 is used to encode the main vehicle information in the target scene information, extract the independent features of the main vehicle from the target scene information, and directly use them as the fused features of the main vehicle. It also encodes the obstacle information in the target scene information, extracting obstacle features, which may specifically include the feature information of any obstacle in the corresponding driving scene. Subsequently, the first attention module obtains the scene representation result based on the correlation between the fused features of the main vehicle and the obstacle features.
[0088] The encoder module 401 can be any available feature encoder, and the first attention module 402 can be a neural network module based on the attention mechanism.
[0089] To further improve the accuracy of scene representation results, in this embodiment, the first scene information may include, in addition to the first vehicle information and the first obstacle information, first traffic instruction information, such as information related to traffic signs like lanes, stop lines, pedestrian crossings, and traffic signs. Specifically, this may include lane width information, stop line location information, pedestrian crossing location information, and semantic information of traffic signs. Similarly, the second scene information may include, in addition to the second vehicle information and the second obstacle information, second traffic instruction information, such as information related to traffic signs like lanes, stop lines, pedestrian crossings, and traffic signs. Specifically, this may include lane width information, stop line location information, pedestrian crossing location information, and semantic information of traffic signs. Based on this, please combine with... Figure 4B In another specific example, the scene representation model 400 may include, in addition to the encoder module 401 and the first attention module 402, a traffic instruction processing module 403 and a second attention module 404.
[0090] The encoder module 401 encodes the main vehicle information in the target scene information, extracting independent features of the main vehicle from the target scene information, and encodes obstacle information in the target scene information, extracting obstacle features from the target scene information, specifically including feature information of any obstacle in the corresponding driving scene. The traffic instruction processing module 403 encodes traffic instruction information in the target scene information, extracting traffic instruction features from the target scene information, specifically including feature information of traffic signs such as lanes, stop lines, pedestrian crossings, and traffic signs in the corresponding driving scene. Subsequently, the second attention module 404 obtains the fused features of the main vehicle based on the correlation between the independent features of the main vehicle and the traffic instruction features. Finally, the first attention module 402 obtains the scene representation result based on the correlation between the fused features of the main vehicle and the obstacle features.
[0091] The encoder module 401 can be any available feature encoder. The first attention module 402 and the second attention module 404 can be neural network modules implemented based on the attention mechanism. Specifically, the first attention module 402 and the second attention module 404 can be two attention modules included in an attention network implemented based on the cross-attention mechanism. The traffic instruction processing module 403 can include multiple feature extraction networks composed of a convolutional neural network (CNN) 4031 and a self-attention module 4032, as well as a feature fusion module 4033. Each feature extraction network corresponds to a traffic instruction sign and is used to encode the traffic instruction information corresponding to the traffic instruction sign to obtain the independent features of the traffic instruction sign. The feature fusion module 4033 is used to fuse the independent features of all traffic instruction signs to obtain traffic instruction features.
[0092] Furthermore, it should be noted that in this embodiment, since the first attention module 402 can be a neural network module implemented based on an attention mechanism, after obtaining the fusion features of the main vehicle and the obstacle features, the first attention module 402's "obtaining scene representation results based on the correlation between the fusion features of the main vehicle and the obstacle features" can include: obtaining query parameters based on the fusion features of the main vehicle, and obtaining key parameters based on the obstacle features; calculating the correlation between the fusion features of the main vehicle and the obstacle features based on the query parameters and the key parameters; and obtaining scene representation results based on the correlation between the fusion features of the main vehicle and the obstacle features. Wherein:
[0093] Query1 = X11 * W1 Q
[0094] Key1 = X12 * W1 K
[0095] Where Query1 is the query parameter, X11 is the fusion feature of the main vehicle, and W... 1Q The first parameter matrix is the learnable matrix, Key1 is the key parameter, X12 is the obstacle feature, and W1 is the obstacle feature. K Let be the learnable second parameter matrix.
[0096] In this embodiment of the disclosure, when the target scene information is the first scene information, the scene representation result is the first representation result; when the target scene information is the second scene information, the scene representation result is the second representation result.
[0097] After obtaining the first representation result and the second representation result, a first loss between the first representation result and the second representation result can be obtained; if the first loss meets the first loss requirement, a trained scene representation model is obtained; if the first loss does not meet the first loss requirement, the parameters of the scene representation model are adjusted, for example, including adjusting the first parameter matrix and the second parameter matrix. The first loss can be a first cosine similarity, and the first loss requirement can be that the first loss is less than a first similarity threshold. The first similarity threshold can be set according to actual application needs, for example, it can be set to 0.3, and this embodiment does not limit this.
[0098] Through the above steps, in this embodiment of the disclosure, target scene information can be input into a scene representation model to extract the fusion features of the main vehicle and obstacle features from the target scene information. The scene representation model then outputs a scene representation result based on the correlation between the fusion features of the main vehicle and the obstacle features. Specifically, when the target scene information is a first scene, the scene representation result is a first representation result; when the target scene information is a second scene, the scene representation result is a second representation result. Finally, the scene representation model is trained based on the first and second representation results. In this process, since the scene representation result is output by the scene representation model based on the correlation between the fusion features of the main vehicle and the obstacle features, the scene representation result can reflect the interaction and influence between feature information, thereby further improving the accuracy of the scene representation result.
[0099] Furthermore, as mentioned earlier, when the scene representation model has such Figure 4B In the case of the network structure shown, "extracting the fusion features of the main vehicle from the target scene information" may include the following steps:
[0100] Extract the unique features of the main vehicle from the target scene information;
[0101] Extract traffic sign features from target scene information;
[0102] Based on the independent features and traffic indication features of the main vehicle, the fused features of the main vehicle are obtained.
[0103] Traffic indication features can include feature information of traffic signs such as lanes, stop lines, pedestrian crossings, and traffic signs in the corresponding driving scenario.
[0104] Through the above steps, in this embodiment of the disclosure, independent features of the main vehicle can be extracted from the target scene information, traffic indication features can be extracted from the target scene information, and based on the independent features of the main vehicle and the traffic indication features, the fused features of the main vehicle can be obtained, so that the fused features of the main vehicle carry not only the independent features of the main vehicle, but also the traffic indication features, thereby enhancing the representativeness of the fused features of the main vehicle and further improving the training effect of the scene representation model.
[0105] In some alternative implementations, the scene representation model training method further includes the following steps:
[0106] Obtain third-scene information for the third driving scenario;
[0107] Obtain information about the fourth driving scenario; there are differences in the distribution of non-critical obstacles between the third and fourth driving scenarios.
[0108] The scene representation model is trained based on the third and fourth scene information.
[0109] The third driving scenario can be the vehicle's historical driving scenario, while the fourth driving scenario can be a derived driving scenario obtained by adjusting the non-critical obstacles in the third driving scenario. Therefore, the distribution of non-critical obstacles differs between the fourth and third driving scenarios, and these non-critical obstacles do not change the driving risk status of the vehicle.
[0110] In this embodiment of the disclosure, the third scene information may include third master vehicle information and third obstacle information. The third master vehicle information may include the master vehicle's speed, acceleration, and pose information in the third driving scene. The third obstacle information may include the obstacle type in the third driving scene, specifically a movable obstacle or a fixed obstacle. If the obstacle type is a movable obstacle, the third obstacle information may also include the obstacle's speed, acceleration, and pose information in the third driving scene. If the obstacle type is a fixed obstacle, the third obstacle information may also include the location of the obstacle within the corresponding area in the third driving scene.
[0111] Similarly, in this embodiment of the disclosure, the fourth scene information may include fourth master vehicle information and fourth obstacle information. The fourth master vehicle information may include the master vehicle's speed, acceleration, and pose information in the fourth driving scene. The fourth obstacle information may include the obstacle type in the fourth driving scene, specifically a movable obstacle or a fixed obstacle. If the obstacle type is a movable obstacle, the fourth obstacle information may also include the obstacle's speed, acceleration, and pose information in the fourth driving scene. If the obstacle type is a fixed obstacle, the fourth obstacle information may also include the location of the obstacle within the corresponding area in the fourth driving scene.
[0112] To further improve the accuracy of scene representation results, in this embodiment of the disclosure, the third scene information may include, in addition to the third vehicle information and the third obstacle information, third traffic instruction information, such as information related to traffic signs such as lanes, stop lines, pedestrian crossings, and traffic signs. Specifically, this may include lane width information, stop line location information, pedestrian crossing location information, and semantic information of traffic signs. Similarly, the fourth scene information may include, in addition to the fourth vehicle information and the fourth obstacle information, fourth traffic instruction information, such as information related to traffic signs such as lanes, stop lines, pedestrian crossings, and traffic signs. Specifically, this may include lane width information, stop line location information, pedestrian crossing location information, and semantic information of traffic signs.
[0113] After obtaining the third scene information of the third driving scenario and the fourth scene information of the fourth driving scenario, the scene representation model can be trained based on the third scene information and the fourth scene information.
[0114] In a specific example, third scene information can be input into the scene representation model to obtain a third representation result output by the scene representation model; fourth scene information can be input into the scene representation model to obtain a fourth representation result output by the scene representation model; and the scene representation model can be trained based on the third and fourth representation results. The "training the scene representation model based on the third and fourth representation results" can include: obtaining a second loss between the third and fourth representation results; obtaining a trained scene representation model if the second loss meets the second loss requirement; and adjusting the parameters of the scene representation model if the second loss does not meet the second loss requirement. The second loss can be a second cosine similarity, and the second loss requirement can be that the second cosine similarity is greater than a second similarity threshold. The second similarity threshold can be set according to actual application needs; for example, it can be set to 0.95, and this embodiment of the present disclosure does not limit this.
[0115] Furthermore, it should be noted that in the embodiments of this disclosure, the scene representation model can be trained first based on the first scene information and the second scene information, and then trained based on the third scene information and the fourth scene information, or it can be trained first based on the third scene information and the fourth scene information, and then trained based on the first scene information and the second scene information. The embodiments of this disclosure do not impose any restrictions on this.
[0116] It should also be noted that, in this embodiment of the present disclosure, the process of training the scene representation model based on the third scene information and the fourth scene information can be referred to the aforementioned description of "training the scene representation model based on the first scene information and the second scene information", which will not be repeated here.
[0117] Through the above steps, in this embodiment of the disclosure, third scene information of a third driving scenario can be obtained; fourth scene information of a fourth driving scenario can be obtained; and a scene representation model can be trained based on the third and fourth scene information. The fourth driving scenario differs from the third driving scenario in the distribution of non-critical obstacles. Therefore, in this embodiment of the disclosure, the scene representation model can also be trained using similar scene pairs (the third and fourth driving scenarios) with different distributions of non-critical obstacles. This can further improve the scene representation model's ability to perceive and process critical obstacles, thereby further increasing the proportion and accuracy of critical obstacle information in the scene representation results, and ultimately improving the accuracy of the scene representation results.
[0118] As mentioned above, in this embodiment of the present disclosure, the third driving scenario can be the historical driving scenario of the main vehicle, and the fourth driving scenario can be a derived driving scenario obtained by adjusting the key obstacles in the third driving scenario. Based on this, in some optional embodiments, the scene representation model training method may further include the following steps:
[0119] From the third driving scenario, select any obstacle whose distance from the main vehicle is greater than the second distance threshold as the second undetermined obstacle;
[0120] If the second posterior decision corresponding to the second undetermined obstacle is to ignore or follow, and the driving path of the second undetermined obstacle in the third driving scenario has no interaction with the driving path of the main vehicle in the third driving scenario, the second undetermined obstacle is determined as a non-critical obstacle; wherein, the second posterior decision is the driving decision made by the main vehicle for the second undetermined obstacle.
[0121] Remove non-critical obstacles in the third driving scenario to obtain the fourth driving scenario.
[0122] The second distance threshold can be set according to actual application requirements. For example, it can be set to 15m. This embodiment does not limit this.
[0123] The fact that the second undetermined obstacle's driving path in the third driving scenario has no interaction with the main vehicle's driving path in the third driving scenario can be understood as: the second undetermined obstacle's driving path in the third driving scenario has no intersection with the main vehicle's driving path in the third driving scenario.
[0124] Please combine Figure 5A and Figure 5B The third driving scenario 501 includes the main vehicle 502. Any obstacle in the third driving scenario 501 whose distance from the main vehicle 502 is greater than a second distance threshold is selected as the second undetermined obstacle 503. The second posterior decision corresponding to the second undetermined obstacle 503 is to ignore it, and the driving path 504 of the second undetermined obstacle 503 in the third driving scenario 501 has no interaction with the driving path 505 of the main vehicle 502 in the third driving scenario 501. Therefore, the second undetermined obstacle 503 can be determined as a non-critical obstacle. Then, the non-critical obstacle is removed in the third driving scenario to obtain the fourth driving scenario 506.
[0125] Through the above steps, in this embodiment of the disclosure, on the one hand, any obstacle in the third driving scenario whose distance from the main vehicle is greater than a second distance threshold can be selected as a second undetermined obstacle. Then, if the second posterior decision corresponding to the second undetermined obstacle is to ignore or follow, and the driving path of the second undetermined obstacle in the third driving scenario has no interaction with the driving path of the main vehicle in the third driving scenario, the second undetermined obstacle is determined as a non-critical obstacle, thereby ensuring the reliability of non-critical obstacles and further improving the training effect of the scene representation model. On the other hand, after determining the second undetermined obstacle as a non-critical obstacle, the non-critical obstacle can be directly removed from the third driving scenario to obtain a fourth driving scenario, thereby simplifying the creation process of the fourth driving scenario and further improving the training efficiency of the scene representation model.
[0126] The following will describe the complete process of a scene representation model training method provided in the embodiments of this disclosure.
[0127] Obtain the first driving scenario.
[0128] Select a target obstacle from the first driving scenario; adjust the target obstacle in the first driving scenario to obtain a pending driving scenario; perform autonomous driving simulation on the master vehicle in the pending driving scenario to obtain the simulated risk state corresponding to the master vehicle driving in the pending driving scenario; if the simulated risk state is different from the historical risk state, identify the target obstacle as a critical obstacle and identify the pending driving scenario as the second driving scenario; wherein, the historical risk state is the actual risk state corresponding to the master vehicle driving in the first driving scenario.
[0129] Obtain the first scene information of the first driving scenario.
[0130] Obtain the second scene information for the second driving scenario.
[0131] Input the first scene information into the scene representation model to obtain the first representation result output by the scene representation model; input the second scene information into the scene representation model to obtain the second representation result output by the scene representation model.
[0132] A first loss is obtained between the first representation result and the second representation result; if the first loss meets the first loss requirement, a trained scene representation model is obtained; if the first loss does not meet the first loss requirement, the parameters of the scene representation model are adjusted. The first loss can be a first cosine similarity, and the first loss requirement can be that the first cosine similarity is less than a first similarity threshold. The first similarity threshold can be set according to actual application needs; for example, it can be set to 0.3. This embodiment of the present disclosure does not limit this.
[0133] Obtain the third driving scenario.
[0134] In the third driving scenario, any obstacle whose distance from the main vehicle is greater than a second distance threshold is selected as the second undetermined obstacle. If the second posterior decision corresponding to the second undetermined obstacle is to ignore or follow, and the driving path of the second undetermined obstacle in the third driving scenario has no interaction with the driving path of the main vehicle in the third driving scenario, the second undetermined obstacle is determined as a non-critical obstacle. The second posterior decision is the driving decision made by the main vehicle regarding the second undetermined obstacle. The non-critical obstacle is removed in the third driving scenario to obtain the fourth driving scenario.
[0135] Obtain the third scene information of the third driving scenario.
[0136] Obtain the fourth scene information for the fourth driving scenario.
[0137] Input the third scene information into the scene representation model to obtain the third representation result output by the scene representation model; input the fourth scene information into the scene representation model to obtain the fourth representation result output by the scene representation model.
[0138] A second loss is obtained between the third and fourth representation results; if the second loss meets the second loss requirement, a trained scene representation model is obtained; if the second loss does not meet the second loss requirement, the parameters of the scene representation model are adjusted. The second loss can be a second cosine similarity, and the second loss requirement can be that the second cosine similarity is greater than a second similarity threshold. The second similarity threshold can be set according to actual application needs; for example, it can be set to 0.95. This embodiment of the present disclosure does not limit this.
[0139] Please see Figure 6 This is a scene diagram illustrating a scene representation model training method provided in an embodiment of this disclosure.
[0140] As described above, the scene representation model training method provided in this disclosure is applied to electronic devices. Electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital processors, servers, blade servers, mainframe computers, in-vehicle systems, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0141] Electronic devices can be used for:
[0142] Obtain the first scene information of the first driving scenario;
[0143] Obtain second scenario information for the second driving scenario; wherein, the distribution of key obstacles differs between the second driving scenario and the first driving scenario, and the key obstacles are used to change the driving risk status of the main vehicle;
[0144] The scene representation model is trained based on the first scene information and the second scene information.
[0145] The first driving scenario can be the historical driving scenario of the main vehicle, and the second driving scenario can be a derived driving scenario obtained by adjusting the key obstacles in the first driving scenario.
[0146] It should be noted that, in the embodiments disclosed herein, Figure 6 The schematic diagrams shown are for illustrative purposes only and are not restrictive. Those skilled in the art can use them as a basis for their own interpretation. Figure 6 The examples may be modified in various obvious ways and / or substitutions, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of this disclosure.
[0147] This disclosure provides an obstacle marking method that can be applied to electronic devices. The following will be combined with... Figure 7 The flowchart shown illustrates an obstacle marking method provided in this disclosure. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order.
[0148] Step S701: Obtain driving scene information for the current driving scenario;
[0149] Step S702: Input the driving scene information into the trained scene representation model to obtain the intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation result based on the driving scene information; wherein, the scene representation model is obtained by training the scene representation model training method.
[0150] Step S703: Based on intermediate parameters, mark the current obstacles in the current driving scene to obtain obstacle marking results.
[0151] The current driving scenario can be the actual driving scenario of the current vehicle.
[0152] The current scene information can include current vehicle information and current obstacle information. The current vehicle information can include the vehicle's speed, acceleration, and pose in the current driving scene. The current obstacle information can include the obstacle type in the current driving scene, specifically, it can be a movable obstacle or a fixed obstacle. If the obstacle type is movable, the current obstacle information can also include the obstacle's speed, acceleration, and pose in the current driving scene. If the obstacle type is fixed, the current obstacle information can also include the obstacle's location within the current driving scene.
[0153] Furthermore, in this embodiment of the disclosure, the scene representation model may have, for example: Figure 4A The network structure shown can also be used to represent scene representation models. Figure 4B The network structure is shown. The scene representation model has, for example, the network structure shown. Figure 4B In the network structure shown, the current scene information can include not only the current vehicle information and the current obstacle information, but also the current traffic instruction information, such as the relevant information of traffic signs such as lanes, stop lines, pedestrian crossings, and traffic signs. Specifically, it can include lane width information, stop line location information, pedestrian crossing location information, and semantic information of traffic signs.
[0154] In this embodiment of the disclosure, the process of "inputting driving scene information into a trained scene representation model to obtain the scene representation model outputting the current scene representation result based on the driving scene information" can be referred to the aforementioned description of "inputting target scene information into a scene representation model to extract the fusion features of the main vehicle from the target scene information and extract obstacle features from the target scene information through the scene representation model; obtaining the scene representation result output by the scene representation model based on the correlation between the fusion features of the main vehicle and the obstacle features", which will not be repeated here.
[0155] After obtaining the obstacle marking results, a behavioral decision can be made for each marked obstacle. This will not be elaborated on here.
[0156] The obstacle marking method provided in this disclosure can acquire driving scene information of the current driving scenario; input the driving scene information into a trained scene representation model to obtain intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation result based on the driving scene information; wherein, the scene representation model is trained by a scene representation model training method; based on the intermediate parameters, the current obstacles in the current driving scenario are marked to obtain obstacle marking results. Since the scene representation model is a learnable neural network model, compared with the prior art, marking the current obstacles in the current driving scenario based on the intermediate parameters obtained by the scene representation model can improve the accuracy of obstacle marking results.
[0157] In this embodiment of the disclosure, the intermediate parameters may include the fused features of the current driver vehicle extracted from the current scene information, and the current obstacle features extracted from the current scene information. Based on this, in some optional implementations, "marking obstacles in the current driving scene based on intermediate parameters" may include the following steps:
[0158] Based on the fusion features of the current main vehicle and the current obstacle features, the criticality of multiple current obstacles in the current driving scene is ranked.
[0159] Select the number of current obstacles with the highest priority from multiple current obstacles as obstacles to be marked;
[0160] Mark the current obstacle to be marked and obtain the obstacle marking results.
[0161] As described above, in this embodiment of the disclosure, the scene representation model may include a first attention module, which may be a neural network module based on an attention mechanism. Therefore, after obtaining the fusion features of the current master vehicle and the current obstacle features, using the fusion features of the current master vehicle and the current obstacle features as intermediate parameters to obtain the scene representation result may include: obtaining the current query parameters based on the fusion features of the current master vehicle, and obtaining the current key parameters based on the current obstacle features; calculating the correlation between the fusion features of the current master vehicle and the current obstacle features based on the current query parameters and the current key parameters; and obtaining the current scene representation result based on the correlation between the fusion features of the current master vehicle and the current obstacle features. Wherein:
[0162] Query2 = X21 * W2 Q
[0163] Key2 = X22 * W2 K
[0164] Where Query2 is the current query parameter, X21 is the fusion feature of the current master vehicle, and W2 is the current query parameter. Q This is the third parameter matrix, which is also the adjusted first parameter matrix. Key2 is the current key parameter, which is also the adjusted second parameter matrix. X22 represents the current obstacle feature, and W2... K This is the fourth parameter matrix.
[0165] After obtaining the current query parameters and current key parameters, the first function can be used to process these parameters to obtain the impact degree of each current obstacle on the vehicle in the current driving scene. Here, a current obstacle is an obstacle perceived by the scene representation model in the current driving scene. This process can be specifically represented as follows:
[0166] attn_score=matmul(Query2,Key2)
[0167] Here, attn_score represents the degree of influence of each current obstacle on the main vehicle in the current driving scene, and matmul() is the first function.
[0168] Subsequently, the second function can be used to sort each current obstacle in the current driving scenario by its approximate criticality. This process can be specifically characterized as follows:
[0169] importance_rank=Sort(attn_score)
[0170] Here, importance_rank is the ranking result of the attention weight of each current obstacle in the current driving scene, and Sort() is the second function.
[0171] Finally, a number of current obstacles with the highest priority ranking can be selected from multiple current obstacles as obstacles to be marked, and these obstacles are then marked to obtain the obstacle marking results. The number of obstacles can be set according to actual application needs; for example, it can be set to 4, but this embodiment does not limit this.
[0172] Through the above steps, in this embodiment of the disclosure, multiple current obstacles in the current driving scene can be ranked by key importance based on the fusion features of the current master vehicle and the current obstacle features. Then, a target number of current obstacles with the highest key importance ranking are selected from the multiple current obstacles as obstacles to be marked, and these obstacles are marked to obtain the obstacle marking result. Since this embodiment of the disclosure ranks multiple current obstacles in the current driving scene by key importance based on the fusion features of the current master vehicle and the current obstacle features, it takes into account the correlation between each current obstacle and the current master vehicle. Therefore, the reliability of the ranking result can be improved, thereby further improving the accuracy of the obstacle marking result.
[0173] The complete process of an obstacle marking method provided in the embodiments of this disclosure will be described below.
[0174] Obtain driving scene information for the current driving scenario.
[0175] Driving scene information is input into a trained scene representation model to obtain intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation result based on the driving scene information. The scene representation model is trained by a scene representation model training method, and the intermediate parameters may include the fusion features of the current master vehicle extracted from the current scene information, and the current obstacle features extracted from the current scene information.
[0176] Based on the fusion features of the current master vehicle and the features of the current obstacles, the keyness of multiple current obstacles in the current driving scene is ranked; a number of current obstacles with the highest keyness ranking are selected from the multiple current obstacles as obstacles to be marked; the obstacles to be marked are marked to obtain the obstacle marking results.
[0177] Please see Figure 8 This is a schematic diagram of a scenario for an obstacle marking method provided in an embodiment of this disclosure.
[0178] As previously described, the obstacle marking method provided in this disclosure is applied to electronic devices. These electronic devices are intended to represent various forms of digital computers, such as in-vehicle systems.
[0179] Electronic devices can be used for:
[0180] Obtain driving scene information for the current driving scenario;
[0181] Driving scene information is input into a trained scene representation model to obtain intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation result based on the driving scene information; wherein, the scene representation model is trained by a scene representation model training method.
[0182] Based on intermediate parameters, the current obstacles in the current driving scene are marked to obtain the obstacle marking results.
[0183] The current driving scenario can be the actual driving scenario of the current vehicle.
[0184] In this embodiment of the disclosure, the driving environment can be perceived through the perception system installed on the vehicle to obtain environmental perception data. Then, the electronic device constructs the current driving scene based on the environmental perception data and obtains the driving scene information of the current driving scene. The perception system may include an imaging unit, lidar, millimeter-wave radar, ultrasonic radar, etc., and this embodiment of the disclosure does not limit the scope of the invention.
[0185] It should be noted that, in the embodiments disclosed herein, Figure 8 The schematic diagrams shown are for illustrative purposes only and are not restrictive. Those skilled in the art can use them as a basis for their own interpretation. Figure 8 The examples may be modified in various obvious ways and / or substitutions, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of this disclosure.
[0186] To better implement the scene representation model training method, this disclosure also provides a scene representation model training device 900, which can be integrated into an electronic device. The following will be combined with... Figure 9 The schematic diagram shown illustrates a scene representation model training device 900 provided in a public embodiment.
[0187] The scene representation model training device 900 may include:
[0188] The first information acquisition unit 901 is used to acquire the first scene information of the first driving scene;
[0189] The second information acquisition unit 902 is used to acquire second scene information of the second driving scene; wherein, the second driving scene and the first driving scene have different distributions of key obstacles, and the key obstacles are used to change the driving risk state of the main vehicle;
[0190] The first model training unit 903 is used to train the scene representation model based on the first scene information and the second scene information.
[0191] In some optional implementations, the first driving scenario is the historical driving scenario of the main vehicle, and the device further includes a first scenario creation unit for:
[0192] Select the target obstacle from the first driving scenario;
[0193] Adjust the target obstacle in the first driving scenario to obtain the undetermined driving scenario;
[0194] Perform autonomous driving simulation on the master vehicle in the undetermined driving scenario to obtain the simulation risk state corresponding to the master vehicle when driving in the undetermined driving scenario;
[0195] When the simulated risk state differs from the historical risk state, the target obstacle is identified as the critical obstacle, and the pending driving scenario is identified as the second driving scenario; the historical risk state is the actual risk state corresponding to the main vehicle driving in the first driving scenario.
[0196] In some optional implementations, the first scene creation unit is used for:
[0197] Determine the scenario category for the first driving scenario;
[0198] In the case of a safe scenario, the first pending obstacle with the first posterior decision to be avoided is selected from the first driving scenario and is taken as the target obstacle; wherein, the first posterior decision is the driving decision made by the main vehicle to the first pending obstacle.
[0199] In the case of a collision scenario or an emergency braking scenario, any obstacle in the first driving scenario whose distance from the main vehicle is less than a first distance threshold is selected as the target obstacle.
[0200] In some optional implementations, the first scene creation unit is used for:
[0201] In the case of a safe scenario, the pose and / or travel time of the target obstacle are adjusted in the first driving scenario to obtain a pending driving scenario.
[0202] And / or, in the case of a collision scenario or an emergency braking scenario, remove the target obstacle in the first driving scenario to obtain a pending driving scenario.
[0203] In some alternative implementations, the first model training unit 903 is used for:
[0204] The target scene information is input into the scene representation model to extract the fusion features of the main vehicle and the obstacle features from the target scene information; wherein, the target scene information is either the first scene information or the second scene information.
[0205] The scene representation model is obtained based on the correlation between the fusion features of the main vehicle and the obstacle features. Wherein, when the target scene information is the first scene information, the scene representation result is the first representation result, and when the target scene information is the second scene information, the scene representation result is the second representation result.
[0206] The scene representation model is trained based on the first and second representation results.
[0207] In some alternative implementations, the first model training unit 903 is used for:
[0208] Extract the unique features of the main vehicle from the target scene information;
[0209] Extract traffic sign features from target scene information;
[0210] Based on the independent features and traffic indication features of the main vehicle, the fused features of the main vehicle are obtained.
[0211] In some optional implementations, the scene representation model training device 900 further includes a second model training unit for:
[0212] Obtain third-scene information for the third driving scenario;
[0213] Obtain information about the fourth driving scenario; there are differences in the distribution of non-critical obstacles between the third and fourth driving scenarios.
[0214] The scene representation model is trained based on the third and fourth scene information.
[0215] In some optional implementations, the third driving scenario is the historical driving scenario of the main vehicle, and the device further includes a second scenario creation unit for:
[0216] From the third driving scenario, select any obstacle whose distance from the main vehicle is greater than the second distance threshold as the second undetermined obstacle;
[0217] If the second posterior decision corresponding to the second undetermined obstacle is to ignore or follow, and the driving path of the second undetermined obstacle in the third driving scenario has no interaction with the driving path of the main vehicle in the third driving scenario, the second undetermined obstacle is determined as a non-critical obstacle; wherein, the second posterior decision is the driving decision made by the main vehicle for the second undetermined obstacle.
[0218] Remove non-critical obstacles in the third driving scenario to obtain the fourth driving scenario.
[0219] In this embodiment of the present disclosure, the specific functions and examples of each unit of the scene representation model training device 900 can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0220] To better implement the obstacle marking method, this disclosure also provides an obstacle marking device, which can be integrated into an electronic device. The following will be combined with... Figure 10The schematic diagram shown illustrates an obstacle marking device 1000 provided in a disclosed embodiment.
[0221] The obstacle marking device 1000 may include:
[0222] The current information acquisition unit 1001 is used to acquire driving scene information of the current driving scene;
[0223] The intermediate parameter acquisition unit 1002 is used to input driving scene information into the trained scene representation model to obtain intermediate parameters obtained by the scene representation model in the process of outputting the current scene representation result based on the driving scene information; wherein, the scene representation model is trained by the method of any one of claims 1 to 8.
[0224] The obstacle marking unit 1003 is used to mark the current obstacles in the current driving scene based on intermediate parameters and obtain the obstacle marking result.
[0225] In some optional implementations, the intermediate parameters include the fused features of the current master vehicle extracted from the current scene information, and the current obstacle features extracted from the current scene information; the obstacle marking unit 1003 is used for:
[0226] Based on the fusion features of the current main vehicle and the current obstacle features, the criticality of multiple current obstacles in the current driving scene is ranked.
[0227] The target number of current obstacles is selected from multiple current obstacles, and the target number of current obstacles with the highest priority ranking are selected as obstacles to be marked;
[0228] Mark the current obstacle to be marked and obtain the obstacle marking results.
[0229] The specific functions and examples of each unit of the obstacle marking device 1000 in this embodiment can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0230] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0231] According to embodiments of this disclosure, this disclosure also provides an electronic device, an autonomous vehicle, a readable storage medium, and a computer program product.
[0232] Figure 11A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. Electronic device 1100 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 1100 may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0233] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. An input / output (I / O) interface 1105 is also connected to bus 1104.
[0234] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0235] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as scene representation model training methods and / or obstacle marking methods. For example, in some embodiments, the scene representation model training methods and / or obstacle marking methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of the scene representation model training method and / or obstacle marking method described above can be performed. Alternatively, in other embodiments, computing unit 1101 can be configured to perform the scene representation model training method and / or obstacle marking method by any other suitable means (e.g., by means of firmware).
[0236] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0237] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0238] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM) or flash memory, optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0239] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0240] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0241] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0242] This disclosure also provides an autonomous driving vehicle, including an electronic device 1100.
[0243] This disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute a scene representation model training method and / or an obstacle marking method.
[0244] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements a scene representation model training method and / or an obstacle marking method.
[0245] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure is achieved, and this is not limited herein. Furthermore, in this disclosure, relational terms such as "first," "second," and "third" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Additionally, in this disclosure, "multiple" can be understood as at least two, and "any" can be understood as any one.
[0246] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A scene representation model training method, comprising: obtaining first scene information of a first driving scene; obtaining second scene information of a second driving scene; wherein the second driving scene and the first driving scene have a distribution difference of a key obstacle, and the key obstacle is used to change a driving risk state of a host vehicle; training the scene representation model based on the first scene information and the second scene information; obtaining third scene information of a third driving scene; obtaining fourth scene information of a fourth driving scene; wherein the third driving scene and the fourth driving scene have a distribution difference of a non-key obstacle; training the scene representation model based on the third scene information and the fourth scene information.
2. The method of claim 1, wherein, The first driving scene is a historical driving scene of the host vehicle, and the method further comprises: selecting a target obstacle from the first driving scene; adjusting the target obstacle in the first driving scene to obtain a pending driving scene; performing automatic driving simulation on the host vehicle in the pending driving scene to obtain a simulation risk state corresponding to the host vehicle driving in the pending driving scene; in the case where the simulation risk state is different from a historical risk state, determining the target obstacle as the key obstacle, and determining the pending driving scene as the second driving scene; wherein the historical risk state is a real risk state corresponding to the host vehicle driving in the first driving scene.
3. The method of claim 2, wherein, The selecting of the target obstacle from the first driving scene comprises: determining a scene category of the first driving scene; in the case where the scene category is a safe scene, selecting a first pending obstacle corresponding to a first posterior decision of avoidance from the first driving scene as the target obstacle; wherein the first posterior decision is a driving decision made by the host vehicle on the first pending obstacle; in the case where the scene category is a collision scene or a sudden braking scene, selecting any obstacle with an interval distance less than a first distance threshold from the host vehicle in the first driving scene as the target obstacle.
4. The method of claim 3, wherein, The adjusting of the target obstacle in the first driving scene to obtain a pending driving scene comprises: in the case where the scene category is a safe scene, adjusting a pose and / or driving time of the target obstacle in the first driving scene to obtain the pending driving scene; and / or, in the case where the scene category is a collision scene or a sudden braking scene, removing the target obstacle in the first driving scene to obtain the pending driving scene.
5. The method according to any one of claims 1 to 4, wherein, The training of the scene representation model based on the first scene information and the second scene information comprises: inputting target scene information into the scene representation model to extract a fusion feature of the host vehicle from the target scene information and extract an obstacle feature from the target scene information through the scene representation model; wherein the target scene information is any one of the first scene information and the second scene information. obtaining a scene characterization result output by the scene characterization model based on a correlation between the fusion feature of the host vehicle and the obstacle feature; wherein, in a case that the target scene information is the first scene information, the scene characterization result is a first characterization result, and in a case that the target scene information is the second scene information, the scene characterization result is a second characterization result; training the scene characterization model based on the first characterization result and the second characterization result.
6. The method of claim 5, wherein, The fusion feature of the host vehicle is extracted from the target scene information, including: extracting an independent feature of the host vehicle from the target scene information; extracting a traffic indication feature from the target scene information; obtaining the fusion feature of the host vehicle based on the independent feature of the host vehicle and the traffic indication feature.
7. The method of claim 1, wherein, The third driving scene is a historical driving scene of the host vehicle, and the method further includes: selecting any obstacle with an interval distance greater than a second distance threshold from the host vehicle in the third driving scene as a second pending obstacle; in a case that a second posterior decision corresponding to the second pending obstacle is to ignore or to follow more closely, and a driving path of the second pending obstacle in the third driving scene does not interact with a driving path of the host vehicle in the third driving scene, determining the second pending obstacle as a non-key obstacle; wherein, the second posterior decision is a driving decision made by the host vehicle on the second pending obstacle; removing the non-key obstacle in the third driving scene to obtain a fourth driving scene.
8. A method for obstacle labeling, comprising: obtaining driving scene information of a current driving scene; inputting the driving scene information into a trained scene characterization model to obtain an intermediate parameter obtained by the scene characterization model in a process of outputting a current scene characterization result based on the driving scene information; wherein, the scene characterization model is trained by the method of any one of claims 1-7; labeling a current obstacle in the current driving scene based on the intermediate parameter to obtain an obstacle labeling result.
9. The method of claim 8, wherein, The intermediate parameter includes a fusion feature of a current host vehicle extracted from the current scene information, and a current obstacle feature extracted from the current scene information; The labeling of the current obstacle in the current driving scene based on the intermediate parameter includes: performing key degree sorting on a plurality of current obstacles in the current driving scene based on the fusion feature of the current host vehicle and the current obstacle feature; selecting a target number of current obstacles with a high key degree from the plurality of current obstacles as pending obstacles; labeling the pending obstacles to obtain the obstacle labeling result.
10. A device for training a scene characterization model, comprising: a first information acquisition unit configured to obtain first scene information of a first driving scene; a second information obtaining unit, configured to obtain second scene information of a second driving scene; wherein the second driving scene is different from the first driving scene in distribution of key obstacles, and the key obstacles are used to change a driving risk state of a host vehicle; a first model training unit, configured to train the scene representation model based on the first scene information and the second scene information; a second model training unit, configured to obtain third scene information of a third driving scene and fourth scene information of a fourth driving scene, and train the scene representation model based on the third scene information and the fourth scene information; wherein the third driving scene is different from the fourth driving scene in distribution of non-key obstacles.
11. An obstacle marking apparatus, comprising: a current information obtaining unit, configured to obtain driving scene information of a current driving scene; an intermediate parameter obtaining unit, configured to input the driving scene information into a trained scene representation model to obtain an intermediate parameter obtained by the scene representation model in a process of outputting a current scene representation result based on the driving scene information; wherein the scene representation model is trained by the method in any one of claims 1 to 7; an obstacle marking unit, configured to mark a current obstacle in the current driving scene based on the intermediate parameter to obtain an obstacle marking result.
12. An electronic device, comprising: at least one processor; a memory connected to the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method in any one of claims 1 to 9.
13. An autonomous vehicle comprising the electronic device of claim 12.
14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, the computer instructions are used to enable the computer to perform the method in any one of claims 1 to 9.
15. A computer program product comprising a computer program which, when executed by a processor, implements the method in any one of claims 1 to 9.