A scene classification method for autonomous driving and apparatus thereof
By constructing a temporal correlation graph to analyze the interaction between dynamic targets and the static environment, the problem of insufficient dynamic element interaction analysis in existing technologies is solved, improving the accuracy and real-time performance of autonomous driving scenario classification and ensuring driving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing autonomous driving scene classification methods are insufficient in handling dynamic element interaction relationships, resulting in a decrease in the accuracy and real-time performance of scene classification, making it difficult to meet high-precision requirements and posing a risk of misjudgment.
By acquiring static element and dynamic target data, a temporal correlation graph is constructed. By combining the motion trajectory and relative motion state of dynamic targets, the interaction relationship between dynamic targets is analyzed in depth, and the scene category judgment is refined.
It improves the accuracy and real-time performance of scene classification, enabling more precise identification of potential risks and ensuring the safety of autonomous driving.
Smart Images

Figure CN120894763B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a scene classification method for automatic driving and a device thereof. BACKGROUND
[0002] With the rapid development of automatic driving technology, scene classification as a core link of environment understanding directly affects the reliability of decision planning. The existing scene classification method for automatic driving often excessively relies on static environmental features such as road markings, fixed buildings, etc. when classifying scenes, and has deficiencies in processing dynamic elements, that is, it fails to deeply analyze the interaction relationship between dynamic targets and the correlation between dynamic targets and static environment. For example, in a crossroad scene, the existing method may only identify the existence of dynamic targets such as vehicles and pedestrians, but cannot effectively judge the avoidance intention between vehicles and pedestrians, the coordination relationship between vehicles and traffic signal lights, etc. The lack of analysis of the interaction of dynamic elements leads to a significant decline in the accuracy and real-time performance of scene classification in complex dynamic scenes, which is difficult to meet the high-precision requirements of automatic driving for scene judgment, and is likely to cause misjudgment risk, affecting the safety of automatic driving. SUMMARY
[0003] The present application provides a scene classification method for automatic driving and a device thereof to meet the high-precision requirements of automatic driving for scene judgment, reduce misjudgment risk, and ensure the safety of automatic driving.
[0004] In a first aspect, the present application provides a scene classification method for automatic driving, comprising:
[0005] Obtaining static element data and dynamic target data in the current scene of an automatic driving vehicle; the static element represents an object in the scene whose position does not change; the dynamic target represents an object in the scene whose position or state changes;
[0006] Extracting and combining features of the static element data based on the type of the static element to obtain static element combined features, and matching the static element combined features with a preset static scene type library to determine an initial scene category;
[0007] Generating motion trajectory data of each dynamic target based on the position coordinates of each dynamic target at different time recorded in the dynamic target data, and analyzing the static element data and the motion trajectory data to construct a time sequence correlation graph of each static element and each dynamic target in the current scene;
[0008] Determining the interaction relationship between dynamic targets based on the relative motion state between dynamic targets identified by the dynamic target data, in combination with the motion speed and motion direction of each dynamic target;
[0009] For each subcategory in the initial scene category, a target scene category of the autonomous vehicle is determined based on the time sequence association graph and the interaction relationship.
[0010] In a second aspect, the present application further provides a scene classification device for autonomous driving, which is applied to the scene classification method for autonomous driving as described in the first aspect; the scene classification device for autonomous driving comprises:
[0011] The acquisition module is configured to acquire static element data and dynamic target data in a current scene of an autonomous vehicle; the static element represents an object in the scene whose position does not change; the dynamic target represents an object in the scene whose position or state changes;
[0012] The initial scene determination module is configured to extract and combine features of the static element data based on types of the static elements, to obtain combined features of the static elements, and to match the combined features of the static elements with a preset static scene type library, to determine an initial scene category;
[0013] The graph construction module is configured to generate motion trajectory data of each dynamic target based on position coordinates of each dynamic target at different time instants recorded in the dynamic target data, and to analyze the static element data and the motion trajectory data, to construct a time sequence association graph of each static element and each dynamic target in the current scene;
[0014] The interaction relationship determination module is configured to determine an interaction relationship between dynamic targets based on relative motion states between the dynamic targets identified based on the dynamic target data, in combination with motion speeds and directions of each dynamic target;
[0015] The scene category determination module is configured to determine a target scene category of the autonomous vehicle based on the time sequence association graph and the interaction relationship for each subcategory in the initial scene category.
[0016] In a third aspect, the present application further provides an electronic device, which comprises a memory configured to store a computer software program, and a processor configured to read and execute the computer software program, to realize the scene classification method for autonomous driving as described in any of the above aspects.
[0017] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium, wherein the storage medium stores a computer software program, and the computer software program is executed by a processor to realize the scene classification method for autonomous driving as described in any of the above aspects.
[0018] In a fifth aspect, the present application further provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the scene classification method for autonomous driving as described in any of the above aspects.
[0019] The scene classification method for automatic driving provided by the embodiment of the present application can construct a time sequence correlation graph by combining motion trajectory data generated according to dynamic target data with static element data, and determine the interaction relationship between dynamic targets based on relative motion state, speed and direction, so as to deeply mine the correlation between dynamic targets and between dynamic targets and static environment, effectively solve the problem of insufficient interaction analysis of dynamic elements in the prior art, in addition, the initial scene category is refined and judged through the time sequence correlation graph and the interaction relationship, which can more comprehensively and accurately capture the dynamic change characteristics of the scene, compared with the prior art classification method which only relies on static characteristics, it can more timely and accurately determine the target scene category, thereby improving the accuracy and real-time performance of scene classification, meeting the high-precision requirements of automatic driving on scene judgment, and more accurately identifying potential risks in the scene, reducing decision-making errors caused by misjudgment of scene classification, and ensuring the safety of automatic driving. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is a flowchart of the scene classification method for automatic driving provided by the embodiment of the present application;
[0021] Figure 2 is a structural schematic diagram of the scene classification device for automatic driving provided by the embodiment of the present application;
[0022] Figure 3 is an embodiment diagram of an electronic device provided by the embodiment of the present application;
[0023] Figure 4 is an embodiment diagram of a computer readable storage medium provided by the embodiment of the present application. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0025] In the description of the present application, the terms "first", "second" are only used for description purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0026] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0027] See Figure 1 , Figure 1 This is a flowchart illustrating the scene classification method for autonomous driving provided by the present invention. In this embodiment, the execution entity of the scene classification method for autonomous driving is a scene classification device. Therefore, the scene classification method for autonomous driving includes:
[0028] Step 10: Obtain static element data and dynamic target data in the current scene of the autonomous vehicle; static elements represent objects whose positions do not change in the scene; dynamic targets represent objects whose positions or states change in the scene.
[0029] Optionally, the scene classification device collects current environmental data of the current scene in which the autonomous vehicle is located using sensors such as LiDAR, cameras, and millimeter-wave radar, and separates static element data and dynamic target data from it. Static elements refer to objects whose positions do not change in a short period of time. Static element data includes parameters (such as shape, position, and color) of objects whose positions remain unchanged, such as road signs (e.g., traffic lights, stop lines), fixed buildings (e.g., streetlights, guardrails), terrain structures (e.g., slope, curve curvature), and obstacles (e.g., bollards, rocks). Dynamic targets refer to objects whose positions or states change over time. Dynamic target data covers real-time information (such as real-time position coordinates, speed, acceleration, and size) of objects whose positions or states change, such as pedestrians, vehicles, and cyclists. It is important to note that after collecting static element data and dynamic target data, the scene classification device preprocesses the data through its built-in preprocessing module. This includes spatiotemporal alignment of the data, unifying the data collected by different sensors into the same time and space coordinate system, with the vehicle's current position as the spatial reference point; removing noise points from the static element data; filtering outliers in the dynamic target data; and prioritizing the static element data and dynamic target data based on the vehicle's driving direction.
[0030] In an embodiment, the scene classification device collects data in an urban road scene, and identifies static element data as: type "traffic signal" (position coordinates (X1, Y1, Z1), size 0.5m x 0.5m x 1.2m), "stop line" (position coordinates (X2, Y2, Z2), length 5m x width 0.3m), "road marking (straight arrow)" (position coordinates (X3, Y3, Z3), length 3m x width 1.5m); dynamic target data as: type "pedestrian" (t1 time position (A1, B1, C1), speed 1.2m / s, direction perpendicular to the vehicle driving direction), "small car" (t1 time position (D1, E1, F1), speed 30km / h, direction same as the vehicle).
[0031] Step 20, feature extraction and combination of static element data based on the type of static element, to obtain static element combination features, and match the static element combination features with the preset static scene type library to determine the initial scene category.
[0032] Optionally, the scene classification device first extracts features by type from the static element data, such as road type, traffic facilities, and surrounding environment, and extracts features of each type, such as road type extraction of urban trunk road, rural road, etc.; traffic facilities extraction of signal light, sign marking, etc.; surrounding environment extraction of commercial area, residential area, etc., and combines (such as splicing combination) the extracted features into static element combination features. In addition, the preset static scene type library in the scene classification device contains feature sets of multiple typical static scenes (such as urban trunk road, highway, rural road, etc.), and then matches the static element combination features with the preset static scene type library, selects the static scene with the highest matching degree in the static scene type library as the initial scene category, and the specific description is as shown in steps 201-205.
[0033] Step 30, based on the position coordinates of each dynamic target at different time recorded in the dynamic target data, generate the motion trajectory data of each dynamic target, and based on the static element data and the motion trajectory data, construct the time sequence correlation graph of each static element and each dynamic target in the current scene.
[0034] Optionally, the scene classification device generates continuous motion trajectory data (such as a polynomial fitting trajectory equation) according to the position coordinates of each dynamic target at different time (such as t1, t2, t3...) in the dynamic target data. At the same time, combined with the static element data (position, type), analyze the space-time relationship between the dynamic target and the static element (such as whether the pedestrian is close to the sidewalk, whether the vehicle is driving along the road marking), and the position change relationship between the dynamic targets, to construct the time sequence correlation graph. The specific description is as shown in steps 301-305.
[0035] Step 40, based on the relative motion state between the dynamic targets identified by the dynamic target data, combined with the motion speed and direction of each dynamic target, determine the interaction relationship between the dynamic targets.
[0036] Optionally, the scene classification device calculates the relative speed (such as the speed difference of two vehicles) and the relative displacement of each dynamic target in the dynamic target data to determine the relative motion state (such as approaching, moving away, parallel, etc.) between the dynamic targets. Then, combined with the motion speed (size) and motion direction (vector) of each dynamic target, analyze the interaction relationship between the targets. For example, if the speed direction of the two dynamic targets is the same and the speed difference is less than a preset threshold (such as 5km / h), it is determined as parallel relationship; if the angle between the speed vectors of the two targets is less than 30° and the distance is continuously decreasing, it is determined that there may be a "car following" relationship; if the direction is perpendicular and the distance is rapidly reduced, it is determined that there may be a "crossing and avoiding" relationship.
[0037] In an embodiment, according to the dynamic target data of steps 10 and 30, the relative motion of the pedestrian and the small car is calculated: the pedestrian speed vector is (0, 1.2, 0) (perpendicular to the vehicle direction), and the small car speed vector is (8.3, 0, 0) (30km / h≈8.3m / s, same direction), the distance between the two is reduced from 15m to 8m at t1-t3 time, and the relative motion state is "approaching each other". Thus, it is determined that the interaction relationship is "the pedestrian and the small car approach each other in the crossing direction, and there is a potential crossing path".
[0038] Step 50, for each subcategory in the initial scene category, determine the target scene category of the autonomous vehicle based on the time sequence association graph and the interaction relationship.
[0039] Optionally, the initial scene category obtained by the scene classification device contains multiple subcategories (such as "urban straight intersection with traffic lights" can be divided into "no conflict intersection when green light" and "pedestrian crossing when red light" subcategories). Therefore, according to the time sequence association graph combined with the interaction relationship between the dynamic targets, the subcategory that best fits the current actual situation is selected from the subcategories of the initial scene category as the target scene category. Specifically, as described in steps 501-505.
[0040] The embodiment of the present application can construct a time sequence correlation graph by combining motion trajectory data generated according to dynamic target data with static element data, and determine the interaction relationship between dynamic targets based on relative motion state, speed and direction, thereby deeply mining the correlation between dynamic targets and between dynamic targets and static environment, effectively solving the problem of insufficient interaction analysis of dynamic elements in the prior art. In addition, the initial scene category is refined by the time sequence correlation graph and the interaction relationship, which can more comprehensively and accurately capture the dynamic change characteristics of the scene. Compared with the existing classification method which only relies on static features, the target scene category can be determined more timely and accurately, thereby improving the accuracy and real-time performance of scene classification, meeting the high-precision requirements of scene judgment of automatic driving, and more accurately identifying potential risks in the scene, reducing decision-making errors caused by misjudgment of scene classification, and ensuring the safety of automatic driving.
[0041] In an embodiment, steps 201-205 are described as follows:
[0042] Step 201: structurally disassembling the static element combination features to obtain multiple independent dimension features; each independent dimension corresponds to the self attribute of a type of static element.
[0043] Optionally, the scene classification device splits the static element combination features according to the category and self attribute of the static element, and each independent dimension corresponds to a specific attribute of a type of static element, such as the width and material of the road element, the number and position of the traffic signal, the type and size of the sign and marking, etc. Through this structured disassembly, complex combination features can be converted into multiple dimension features that can be analyzed independently, laying a foundation for subsequent matching work. And reduce the complexity of subsequent feature matching, make the subsequent matching more targeted, and avoid interference between different dimension features.
[0044] In an embodiment, taking the static element combination feature "containing traffic signal, stop line, straight arrow marking, road width 8m" as an example, the scene classification device performs structured disassembly. Among them, the dimension feature corresponding to the traffic signal is "there is a traffic signal (1 in number)"; the dimension feature corresponding to the stop line is "there is a stop line (length 5m x width 0.3m)"; the dimension feature corresponding to the straight arrow marking is "there is a straight arrow marking (length 3m x width 1.5m)"; the dimension feature corresponding to the road is "road width 8m". Multiple independent dimension features are obtained after disassembly.
[0045] Step 202: extracting the feature dimension information of the static elements contained in each scene type in the static scene type library, and constructing a scene type and dimension index table.
[0046] Optionally, the scene classification device includes city main road, highway, rural road, school area and the like in the static scene type library, traverses each scene type in the static scene type library, extracts the feature dimension information corresponding to the attributes of all static elements contained by each scene type, and then corresponds the scene type to the feature dimension information to form a scene type and dimension index table. The index table can clearly show the feature dimensions possessed by different scene types, facilitating subsequent matching degree calculation.
[0047] In an embodiment, the static scene type library includes scene types such as "city straight intersection with traffic light" and "rural T-shaped intersection without traffic light". For the "city straight intersection with traffic light", the extracted feature dimension information includes "existence of traffic light (number ≥ 1)", "existence of stop line (length ≥ 4 m)", "existence of straight arrow marking (length ≥ 2 m)", "road width 6-10 m" and the like; for the "rural T-shaped intersection without traffic light", the extracted feature dimension information includes "non-existence of traffic light", "possible existence of stop line (length ≤ 3 m)", "existence of T-shaped intersection marking", "road width 3-5 m" and the like. The scene classification device arranges these information to construct a scene type and dimension index table.
[0048] In step 203, the matching degree of each dimension feature and the dimension index of each scene type is calculated based on the scene type and dimension index table to construct a dimension matching degree matrix. Each element in the dimension matching degree matrix reflects the degree of fit of a certain feature dimension and a certain scene type in the dimension layer.
[0049] Optionally, the scene classification device calculates the matching degree between each dimension feature obtained in step 201 and the corresponding dimension index of each scene type in the index table one by one based on the scene type and dimension index table. The calculation of the matching degree can be performed according to the specific attributes of the features by using corresponding methods, such as calculating the difference ratio for numerical attributes and directly judging whether they are consistent for yes / no attributes. Then, these matching degree values are arranged according to the corresponding relationship between the dimension features and the scene types to form a dimension matching degree matrix. The rows of the matrix represent sub-dimension features, the columns represent scene types, and the elements are matching degree values.
[0050] In an embodiment, taking the dimension feature of step 201 and the scene type and dimension index table of step 202 as an example. For the dimension feature of "existence of traffic signal (number 1)", the matching degree with the dimension index of "existence of traffic signal (number > 1)" in the "straight intersection of urban city with signal" is 1.0; the matching degree with the dimension index of "no traffic signal" in the "T intersection of rural area without signal" is 0. For the dimension feature of "road width 8m", the matching degree with the dimension index of "road width 6-10m" in the "straight intersection of urban city with signal" is determined as 0.9 by calculating (8-6) / (10-6)=0.5 and combining the degree of compliance in the range; the matching degree with the dimension index of "road width 3-5m" in the "T intersection of rural area without signal" is 0. After the matching degrees of all dimension features and each scene type are calculated, the dimension matching degree matrix is constructed.
[0051] Step 204, based on the dimension matching degree matrix, the scene types meeting the preset dimension matching threshold are screened out to obtain the candidate scene types.
[0052] Optionally, the scene classification device first sets a preset dimension matching threshold, which is determined according to the actual application scene and the classification accuracy requirement. Then, the dimension matching degree matrix is checked, for each scene type, the comprehensive value (such as the average value, the minimum value, etc.) of the matching degrees of all dimension indexes and corresponding dimension features is calculated, when the comprehensive value is greater than or equal to the preset dimension matching threshold, the scene type is screened out as the candidate scene type.
[0053] In an embodiment, taking the preset dimension matching threshold of 0.7 as an example. For the "straight intersection of urban city with signal", the average value of the dimension matching degrees is calculated, which is assumed to be 0.85, greater than 0.7, and is screened as the candidate scene type; for the "T intersection of rural area without signal", the average value of the dimension matching degrees is 0.3, less than 0.7, and is not screened. After screening, the candidate scene types include the "straight intersection of urban city with signal" and the like.
[0054] Step 205, matching and sorting are performed for each scene type in the candidate scene types to determine the initial scene category.
[0055] Optionally, the scene classification device further matches and sorts the detail features of each scene type in the candidate scene types with the static scene type library to finally obtain the initial scene category, which is specifically described in steps 2051-2054.
[0056] The embodiment of the application determines the initial scene category efficiently and accurately through fine processing of static element combination features and accurate matching with a static scene type library. The complex features are first decomposed into independent dimensions, and then the range is gradually narrowed through steps such as index table construction and matching degree matrix, and finally the most suitable initial scene category is screened out, thereby providing a reliable basis for more accurate scene classification combined with dynamic information.
[0057] In an embodiment, steps 2051-2054 are described as follows:
[0058] In step 2051, the feature details of each candidate scene type are extracted from the static scene type library for each scene type in the candidate scene type, to obtain scene type detail features.
[0059] Optionally, the scene classification device retrieves the complete feature information corresponding to each scene type in the candidate scene type from the static scene type library, which includes not only the basic dimensional feature requirements, but also the specific parameter range of each dimensional feature, the association relationship between features, and other details, to form scene type detail features. Through extraction of these detail features, a basis is provided for subsequent accurate comparison of each dimension.
[0060] In an embodiment, the candidate scene types are "urban straight intersection with traffic lights" and "urban left-turn intersection with traffic lights". The scene classification device extracts the detail features of "urban straight intersection with traffic lights" from the static scene type library: traffic lights (1-2 in number, located directly above the intersection), stop line (length 5-6 m, width 0.3-0.4 m, distance from the intersection 3-5 m), straight arrow marking (length 3-4 m, width 1.5-2 m, arrow pointing in the same direction as the road), road width (6-10 m, four lanes in both directions); and extracts the detail features of "urban left-turn intersection with traffic lights": traffic lights (2-3 in number, including left-turn arrow light), stop line (length 5-6 m), left-turn arrow marking (length 3-4 m), road width (6-10 m), etc.
[0061] In step 2052, the dimensional features of each dimension are compared with the scene type detail features in each dimension, to obtain the feature detail matching degree of each candidate scene type in each independent dimension.
[0062] Optionally, the scene classification device compares the dimensional features obtained in step 201 with the scene type detail features of each candidate scene type extracted in step 2051 in each dimension. Different comparison methods are used for different types of dimensional features, such as calculating the degree of fit between the actual value and the standard range for numerical features, and judging whether the attribute features are completely matched, to obtain the feature detail matching degree of each candidate scene type in each independent dimension.
[0063] In an embodiment, taking the dimensional features of step 201 (traffic signal: 1; stop line: 5m x 0.3m; straight arrow marking: 3m x 1.5m; road width 8m) as an example, compare them with the detailed features of "urban straight intersection with signal": the number of traffic signals is in the range of 1-2, with a matching degree of 1.0; the length of the stop line is 5m, which is in the range of 5-6m, and the width is 0.3m, which is consistent with 0.3-0.4m, with a matching degree of 0.9; the size of the straight arrow marking is within the standard range, with a matching degree of 1.0; the road width is 8m, which is in the range of 6-10m, with a matching degree of 0.9. Compared with "urban left turn intersection with signal": there is no left turn arrow marking, so the matching degree of this dimension is 0; the number of traffic signals is consistent, with a matching degree of 1.0; the matching degrees of the stop line and the road width are 0.9 and 0.9 respectively.
[0064] Step 2053: integrate the matching degrees of the feature details of each candidate scene type to obtain the total matching degree of the feature details of each candidate scene type.
[0065] Optionally, the scene classification device integrates the matching degrees of the feature details of each candidate scene type in each independent dimension, taking into account the differences in importance of different dimensions in scene classification, and finally obtains a value that can comprehensively reflect the overall matching degree of the candidate scene type with the current static element combination feature, i.e. the total matching degree of the feature details.
[0066] Specifically, after obtaining the matching degrees of each dimension of each candidate scene type, the scene classification device integrates these dimension matching degrees by assigning weights according to the overall importance of different dimensions in scene classification. For example, the road structure dimension and the road marking dimension have a greater impact on the determination of the scene category, and the weights are set to 0.4 and 0.3 respectively, and the weight of the fixed facility dimension is set to 0.3. The total matching degree is the sum of the product of each dimension matching degree and the corresponding weight, and the calculation formula is total matching degree = (road structure dimension matching degree x 0.4) + (road marking dimension matching degree x 0.3) + (fixed facility dimension matching degree x 0.3). If there are other dimensions, they are also included in the calculation in the same way, ensuring that the sum of the weights of all dimensions is 1.
[0067] In an embodiment, taking the "urban straight intersection with traffic lights" as an example, the matching degree of each dimension is 1.0 (traffic lights), 0.9 (stop line), 1.0 (straight arrow marking), and 0.9 (road width). In this scenario, traffic lights, straight arrow marking, and stop line have a greater impact, and the determination weights are all 0.3, and the weight of road width is 0.1, so the total matching degree is 1.0*0.3+0.9*0.3+1.0*0.3+0.9*0.1=0.96. For the "urban left-turn intersection with traffic lights", the matching degree of each dimension is 1.0 (traffic lights), 0.9 (stop line), 0 (left-turn arrow marking), and 0.9 (road width), and the total matching degree after integration is 1.0*0.3+0.9*0.3+0*0.3+0.9*0.1=0.66.
[0068] In step 2054, each scene type in the candidate scene types is prioritized based on the total matching degree of the feature details from high to low, and the scene type at the top of the ranking is determined as the initial scene category.
[0069] Optionally, the scene classification device prioritizes the candidate scene types in the order of the total matching degree of the feature details from high to low. After the ranking is completed, the candidate scene type at the top of the ranking is determined as the initial scene category, which can most accurately reflect the static features of the current scene.
[0070] In an embodiment, the ranking result of the candidate scene types is "urban straight intersection with traffic lights" (0.96) and "urban left-turn intersection with traffic lights" (0.66). The scene classification device determines the "urban straight intersection with traffic lights" at the top of the ranking as the initial scene category.
[0071] The embodiment of the application realizes the goal of accurately selecting the initial scene category from the candidate scene types by fine feature extraction, dimension-by-dimension comparison, and matching degree integration, can effectively distinguish similar scene types (such as straight intersection and left-turn intersection), greatly improves the accuracy of the initial scene category, and provides a solid foundation for subsequent scene classification combined with dynamic information.
[0072] In an embodiment, steps 301-305 are described as follows:
[0073] In step 301, based on the position coordinates of the static elements in the static element data and the motion trajectory data of the dynamic targets, the spatial belonging relationship of each dynamic target with the static elements at each time is determined, and an initial association pair is constructed based on the spatial belonging relationship.
[0074] Optionally, the scene classification device determines the spatial range of the static element according to the position coordinates of the static element (such as the coverage area of the traffic light, the boundary range of the zebra crossing, etc.), and then compares the position coordinates of the dynamic target at each time with the spatial range to determine whether the dynamic target is in the spatial range of a certain static element or is affected by the spatial range, thereby determining the spatial attribution relationship, and then combining the static element and the dynamic target that exist in the attribution relationship into an initial association pair.
[0075] In an embodiment, among the static elements, the spatial range of the zebra crossing is a rectangular area with coordinates (X10, Y10) to (X11, Y11), and the spatial influence range of the traffic light is the intersection area (X12, Y12) to (X13, Y13) controlled by the traffic light. The dynamic target pedestrian A is in the zebra crossing range at time t2 with position (X6, Y6), and is in the intersection area controlled by the traffic light at time t3 with position (X9, Y9), so it is determined that the pedestrian A has a spatial attribution relationship with the zebra crossing at time t2, and has a spatial attribution relationship with the traffic light at time t3, forming initial association pairs (pedestrian A, zebra crossing, t2) and (pedestrian A, traffic light, t3). The vehicle B is in the intersection area controlled by the traffic light at time t3 with position (X10, Y10), forming an initial association pair (vehicle B, traffic light, t3).
[0076] Step 302, based on the initial association pair, extracting the instantaneous feature in the motion trajectory data of the dynamic target in each association pair and the interaction feature of the original attribute of the static element itself, obtaining the association interaction feature value.
[0077] Optionally, the scene classification device extracts the instantaneous feature from the motion trajectory data of the dynamic target for each initial association pair, such as the motion speed, acceleration, motion direction angle, etc. at that time; at the same time, the original attribute of the static element itself is obtained, such as the current color of the traffic light, the indication direction of the road sign, the height of the guardrail, etc. By analyzing the interaction between the instantaneous feature and the original attribute, the association interaction feature value is obtained, which quantifies the interaction degree of the two.
[0078] In one embodiment, for the initial association pair (vehicle B, traffic light, t3), the instantaneous characteristics of the dynamic target vehicle B at time t3 are a speed of 30 km / h and a direction angle of 90° (northward); the inherent attributes of the static element traffic light are a red light and an installation height of 3m. Analysis shows that the vehicle should slow down and stop when the light is red, but the vehicle's current speed is relatively high. The interaction between the two is that the vehicle does not act according to the traffic light instruction, resulting in an association interaction feature value of 0.3 (assuming 0 fully conforms to the interaction rules and 1 does not). For (pedestrian A, zebra crossing, t2), pedestrian A's instantaneous speed is 1.2 m / s and the direction angle is 0° (eastward). The zebra crossing attributes are for pedestrians to cross the street and a length of 5m. The pedestrian walks normally along the zebra crossing direction, resulting in an interaction feature value of 0.9 (indicating that the height conforms to the interaction rules).
[0079] Step 303: Calculate the coupling degree of motion features of different dynamic targets under the association of the same static element based on the associated interaction feature value.
[0080] Optionally, the scene classification device collects the associated interaction feature values of all dynamic targets associated with the same static element. By analyzing the similarity and correlation of the motion features such as speed changes and direction adjustments of these dynamic targets during the motion process, the motion feature coupling degree between different dynamic targets is calculated. This motion feature coupling degree reflects the degree of coordination of the motion features of each dynamic target under the influence of the same static element.
[0081] Specifically, the motion feature coupling degree can be calculated as follows:
[0082] Let O = {o1, o2, ..., o} be the set of dynamic targets associated with the same static element. n}, each dynamic target o i The motion feature vector is F i =(v i1 ,v i2 ,…,v ik ), where v ij Let j be the j-th motion feature parameter (e.g., velocity, direction angle). The coupling degree is calculated using the maximum mean difference (MMD) based on the kernel function.
[0083]
[0084] Where φ is the kernel function mapping, which maps the feature vector to the high-dimensional regenerating kernel Hilbert space. The smaller the MMD(O) value, the higher the coupling degree of motion features between dynamic targets.
[0085] In an embodiment, the dynamic targets associated with the same static element traffic light are pedestrian A and vehicle B. The motion characteristics of pedestrian A at time t3 are speed 1.0 m / s and direction east, and the motion characteristics of vehicle B at time t3 are speed 30 km / h and direction north. Through analysis, under the influence of the traffic light, the pedestrian crosses the street at normal speed and the vehicle does not slow down, and the motion characteristics have low coordination, and the statistical motion characteristic coupling degree is 0.2 (0 represents no coupling and 1 represents high coupling).
[0086] In step 304, the transfer strategy of the dynamic target between different static elements is analyzed based on the initial association pair, and a time sequence constraint relationship of the static element to the dynamic target is constructed.
[0087] Optionally, the scene classification device observes the change of the association pair formed by the dynamic target and different static elements at different times, analyzes the transfer strategy adopted by the dynamic target when transferring from association with one static element to association with another static element, such as path selection, speed change, and whether to follow the spatial boundary of the static element, and the like. According to the transfer strategy, a time sequence constraint relationship of the static element to the dynamic target is constructed, that is, the time and space restrictions that the dynamic target needs to follow when transferring between different static elements.
[0088] In an embodiment, the initial association pairs of the dynamic target pedestrian A are (pedestrian A, zebra crossing, t2) and (pedestrian A, traffic light, t3) in sequence. When transferring from association with the zebra crossing to association with the traffic light, the transfer strategy is to walk along the zebra crossing in a straight line and keep the speed at about 1.2 m / s, without exceeding the spatial range of the zebra crossing and the intersection. Thus, the time sequence constraint relationship is constructed: the pedestrian A needs to transfer from the zebra crossing area to the intersection area controlled by the traffic light at a speed of 1-1.5 m / s within the time period from t2 to t3, that is, the time sequence constraint relationship of the static elements zebra crossing and traffic light to the pedestrian A is “uniform speed transfer along the specified path”.
[0089] In step 305, a time sequence association graph is constructed based on the time sequence constraint relationship, the motion characteristic coupling degree, and the association interaction feature value.
[0090] Optionally, the scene classification device integrates the time sequence constraint relationship (including the time sequence influence of the static element attribute) of step 304, the motion characteristic coupling degree of step 303, and the association interaction feature value of step 302 to construct a time sequence association graph, which is specifically described in steps 3051-3054.
[0091] The embodiment of the application can accurately and dynamically reflect the complex relationship between each element and the dynamic target in time and space in the scene by analyzing the spatial attribution, interactive characteristics, motion coupling and timing constraints of the dynamic target and static elements, providing comprehensive and detailed timing correlation basis for subsequent analysis of the interaction between dynamic targets and determination of the target scene category, so that the automatic driving system can more deeply understand the dynamic change law of the current scene.
[0092] In an embodiment, steps 3051-3054 are described as follows:
[0093] In step 3051, constraint analysis is performed based on the timing constraint relationship and the associated interactive feature value to obtain timing dependency features of the dynamic target in the static element sequence; the timing dependency features reflect the response of the dynamic target behavior to the change of the static element; the static element sequence refers to a continuous sequence formed by the static elements associated with the dynamic target at different time instants in time order.
[0094] Optionally, the scene classification device analyzes the motion behavior change of the dynamic target in the static element sequence composed of different static elements, judges whether the behavior of the dynamic target responds to the attribute change of the static element, and thus obtains the timing dependency features according to the constructed timing constraint relationship and the associated interactive feature value. The features can reflect the degree of dependence and response mode of the dynamic target on the change of the static element in the time sequence.
[0095] In an embodiment, the static element sequence is (zebra crossing, t2)→(traffic signal, t3), and the timing constraint relationship of the dynamic target pedestrian A is uniform transfer along the zebra crossing. The associated interactive feature value is 0.9 when associated with the zebra crossing, and when associated with the traffic signal, if the signal is green, the pedestrian speed is maintained at 1.2 m / s (responding to the green light that can pass), and the feature value is 0.8; if it becomes red, the pedestrian slows down to 0.5 m / s (responding to the red light needs to be cautious), and the feature value is 0.6. Thus, the timing dependency features are obtained: the motion speed of the pedestrian A adjusts with the change of the static element (traffic signal light color), showing strong timing dependency.
[0096] In step 3052, constraint screening is performed based on the motion feature coupling degree and the timing dependency features to construct a multi-element interactive constraint network containing static elements and dynamic targets.
[0097] Optionally, the scene classification device first performs constraint screening according to the motion feature coupling degree and the timing dependency features, eliminates the association relationship with extremely low coupling degree and not obvious timing dependency features, and retains the element association with significant interaction. Then, the static elements and the dynamic targets are taken as network nodes, and the connection between the nodes represents the interactive constraint relationship between them after screening, and a multi-element interactive constraint network is constructed, which can clearly show the complex interaction between elements.
[0098] In an embodiment, the interaction between pedestrian A and vehicle B under the association of traffic signal light with motion feature coupling degree of 0.2, and the interaction between pedestrian A and zebra crossing and traffic signal light with time sequence dependent feature strength, and the interaction between vehicle B and traffic signal light are reserved. In the constructed multi-element interaction constraint network, the nodes include pedestrian A, vehicle B, zebra crossing and traffic signal light; the constraint relationship of the edges is: pedestrian A-zebra crossing (walking along it), pedestrian A-traffic signal light (adjusting speed according to light color), vehicle B-traffic signal light (driving according to light color), and pedestrian A-vehicle B (crossing and avoiding).
[0099] In step 3053, the multi-element interaction constraint network and the motion feature coupling degree are combined to obtain the association strength value of the edge in the graph.
[0100] Optionally, the scene classification device combines the multi-element interaction constraint network and the motion feature coupling degree, and calculates the association strength value of each edge in the network through an association strength preset function. The value comprehensively considers the close degree of the interaction constraint relationship between the elements and the coupling degree of the motion feature, and the greater the value, the stronger the association.
[0101] Specifically, when calculating the association strength value of each edge in the network, the association strength preset function is as follows:
[0102]
[0103] Wherein, S e represents the association strength value; W e represents the constraint relationship weight, which is used to describe the importance of the constraint relationship of the edge e in the multi-element interaction constraint network, and the value range is [0, 1], which is pre-set by the static element attribute and the traffic rule, such as the constraint relationship between the pedestrian and the zebra crossing (walking along it), which conforms to the basic traffic rule, W e = 0.9; the constraint relationship between the vehicle and the traffic signal light (driving according to the light color), which involves the core safety rule, W e = 0.95; C e represents the motion feature coupling degree; H e represents the space-time interaction entropy, which is used to quantify the interaction uncertainty of the element pair corresponding to the edge e in the time sequence and the space range; P(t) represents the probability of the interaction of the static element pair at time t (calculated according to the initial association pair frequency in step 301); the interaction uncertainty is converted into a deterministic weight (the lower the entropy value, the higher the determinacy) through (1-H e ), which is multiplied by the constraint relationship weight W e , and then the motion feature coupling degree C e is taken as the average value, so as to realize the comprehensive reflection of the association strength.
[0104] In an embodiment, in the multi-element interaction constraint network, the interaction constraint relationship between pedestrian A and the zebra crossing is close, and the motion feature coupling degree is expressed as that the pedestrian completely follows the zebra crossing attribute in the association, and the association strength value obtained after combination is 0.8; in the interaction between vehicle B and the traffic signal lamp, the vehicle does not follow the red light attribute at present, but the constraint relationship itself is important, and the association strength value obtained by combining the motion feature coupling degree is 0.5; the cross-avoidance constraint relationship between pedestrian A and vehicle B is critical, the motion feature coupling degree is 0.2, and the association strength value is 0.6.
[0105] In step 3054, the association relationships of all time moments are integrated based on the association strength value and the initial association, and a time sequence association graph containing static elements, dynamic targets and time sequence associations thereof is generated.
[0106] Optionally, the scene classification device integrates the association relationships of all time moments according to the association strength value and the initial association, and for the association relationship with the association strength value reaching a certain threshold value, the association relationship is reserved in the form of an edge in the graph, and the association strength value and the corresponding time moment are marked. Finally, the time sequence association graph contains static element and dynamic target nodes, and association edges between the nodes with strength information changing over time, which can comprehensively reflect the association of each element in the time sequence.
[0107] In an embodiment, the association strength value threshold value is set to 0.4, and the association relationship with the association strength value greater than or equal to 0.4 is reserved. After integration, the nodes in the time sequence association graph are pedestrian A, vehicle B, zebra crossing and traffic signal lamp; the edges include (pedestrian A-zebra crossing, t2, 0.8), (pedestrian A-traffic signal lamp, t3, 0.7), (vehicle B-traffic signal lamp, t3, 0.5), (pedestrian A-vehicle B, t4, 0.6), and the association and strength of each element at different time moments are completely presented.
[0108] The embodiment of the application realizes the spatio-temporal association visualization of static elements and dynamic targets through the constructed time sequence association graph, retains the logical constraint relationship of “signal lamp-vehicle-pedestrian”, and quantifies the closeness of interaction through the association strength value. The embodiment of the application can accurately capture the dynamic evolution of the scene (such as the behavior switching from red light to green light) and the potential risk (such as the low-intensity associated pedestrian who may suddenly rush in), provides a multi-dimensional and time-sequenced analysis basis for the subsequent judgment of the target scene category, and solves the limitation of the traditional method which only depends on static features or isolated dynamic data.
[0109] In an embodiment, steps 501-505 are described as follows:
[0110] In step 501, for each subcategory in the initial scene category, a dynamic constraint domain is constructed based on the spatio-temporal association relationship between the static elements and the dynamic targets in the time sequence association graph; the dynamic constraint domain limits the motion range boundary of the dynamic target and the association dimension of the static element under each subcategory.
[0111] Optionally, the scene classification device extracts the spatio-temporal association relationship between static elements and dynamic targets from the time-series association graph for each subcategory of the initial scene category, including the position change of dynamic targets at different time instants relative to static elements, the spatial overlap degree of motion trajectories and static elements, etc. Based on these relationships, a dynamic constraint domain is constructed for each subcategory, which clearly specifies the motion range boundary of dynamic targets (such as vehicles must not exceed lane lines, pedestrians must not enter motor vehicle lanes, etc.) and the association dimension of static elements (such as the association dimension of traffic lights and vehicles is the light color state and vehicle driving state, and the association dimension of zebra crossings and pedestrians is the position coverage and crossing behavior) in the subcategory scene.
[0112] In an embodiment, taking the subcategories of the initial scene category "urban intersection" including "urban intersection with pedestrian and vehicle crossing avoidance" and "urban intersection with only vehicle traffic" as examples. For the "urban intersection with pedestrian and vehicle crossing avoidance" subcategory, based on the time-series association graph, pedestrians mainly move within the zebra crossing and intersection range, and vehicles drive within the lane and need to avoid pedestrians. The constructed dynamic constraint domain is: the motion range boundary of dynamic target pedestrians is the area within 1 meter on both sides of the zebra crossing and intersection, and the association dimension with static elements includes the coverage relationship between pedestrian position and zebra crossing, and the light color association between pedestrian motion direction and traffic light; the motion range boundary of dynamic target vehicles is within the lane lines, and the association dimension with static elements includes the distance between vehicle position and lane line, and the light color association between vehicle speed and traffic light.
[0113] Step 502, based on the dynamic constraint domain and the interaction relationship between dynamic targets, and combining the matching degree of each interaction relationship within the dynamic constraint domain, determine the constraint-adapted interaction that completely falls within the dynamic constraint domain.
[0114] Optionally, the scene classification device compares the interaction relationship between dynamic targets with the dynamic constraint domain of each subcategory, analyzes the matching degree of each interaction relationship within the dynamic constraint domain, i.e., judges whether the interaction relationship conforms to the limitation of the motion range boundary and the association dimension in the dynamic constraint domain. The interaction relationship that completely falls within the dynamic constraint domain is determined as the constraint-adapted interaction; if the interaction relationship exceeds the motion range boundary or does not conform to the association dimension requirement, it is not a constraint-adapted interaction.
[0115] In an embodiment, taking the interaction between dynamic targets as an example of "pedestrian A and vehicle B cross-avoiding", for the "urban intersection with pedestrian and vehicle cross-avoiding" subclass, the dynamic constraint domain limits the pedestrian to move within 1 meter on both sides of the zebra crossing and the intersection, and the vehicle to drive within the lane. After comparison, the motion range of pedestrian A is within 0.5 meters on both sides of the zebra crossing and the intersection, and vehicle B drives within its own lane, which fully meets the limitation of the dynamic constraint domain, the matching degree is 100%, and it is determined as a constraint-adapted interaction. For the "urban intersection with only vehicle passing" subclass, the dynamic constraint domain limits no pedestrian to enter the intersection area, and in this interaction, pedestrian A enters the intersection, which exceeds the boundary of the motion range, the matching degree is 0, and it does not belong to the constraint-adapted interaction.
[0116] In step 503, the time dimension of the constraint-adapted interaction and the time sequence correlation graph is concatenated with the interaction between dynamic targets in chronological order to form a time sequence interaction event chain under each subclass. Each interaction event in the time sequence interaction event chain includes start and end time, specific location of interaction occurrence, and behavior result.
[0117] Optionally, the scene classification device concatenates the constraint-adapted interactions under each subclass in chronological order, and at the same time, integrates the time dimension information of the time sequence correlation graph to form a time sequence interaction event chain. Each interaction event includes start and end time (such as t1 to t2), specific location of interaction occurrence (such as intersection center coordinates (X1, Y1)), and behavior result (such as pedestrian successfully crossing the street, vehicle slowing down to avoid, etc.), which completely presents the time development context of the interaction between dynamic targets under this subclass.
[0118] In an embodiment, the constraint-adapted interactions under the "urban intersection with pedestrian and vehicle cross-avoiding" subclass are in chronological order: at time t2, pedestrian A enters the zebra crossing (position (X5, Y5)), at time t3, vehicle B approaches the intersection (position (X10, Y10)), at time t4, pedestrian A and vehicle B cross-avoid in the intersection center (position (X11, Y11)), at time t5, pedestrian A leaves the intersection (position (X12, Y12)), and vehicle B continues to drive. Concatenate these interactions to form a time sequence interaction event chain: event 1 (t2-t3, (X5, Y5) to (X10, Y10), pedestrian enters zebra crossing, vehicle approaches intersection) → event 2 (t3-t4, (X10, Y10) to (X11, Y11), vehicle slows down, pedestrian continues to cross the street) → event 3 (t4-t5, (X11, Y11) to (X12, Y12), pedestrian leaves the intersection, vehicle accelerates to drive).
[0119] In step 504, the trigger correlation between dynamic targets and static elements when each interaction event occurs is extracted for the time sequence interaction event chain to obtain static trigger features.
[0120] Optionally, the scene classification device analyzes the triggering relationship between the dynamic target and the static element when the event occurs, i.e., whether the attribute or state change of the static element triggers the behavior change of the dynamic target, for each interaction event in the time sequence interaction event chain, so as to extract the static triggering feature. For example, the traffic signal light triggers the vehicle to slow down when the traffic signal light changes from green to red, and the existence of the zebra crossing triggers the pedestrian to cross the street. These are typical static triggering features, which enable the feature to reflect the driving effect of the static element on the interaction behavior of the dynamic target.
[0121] In an embodiment, in the time sequence interaction event chain of the "urban intersection with pedestrian and vehicle cross avoidance" subcategory, the traffic signal light changes to green at time t2 in event 1, triggering pedestrian A to start crossing the zebra crossing, at which time the color change of the static element traffic signal light forms a triggering association with the crossing behavior of the dynamic target pedestrian A; vehicle B detects the zebra crossing and pedestrian A at time t3 in event 2, triggering vehicle B to slow down, and the position information of the static element zebra crossing forms a triggering association with the deceleration behavior of the dynamic target vehicle B. The extracted static triggering features include "traffic signal light green → pedestrian crossing start" and "zebra crossing exists → vehicle slows down to avoid".
[0122] Step 505, determining the target scene category based on the static triggering feature and the time sequence interaction event chain.
[0123] Optionally, the scene classification device finally determines the target scene category according to the determined static triggering feature and the time sequence interaction event chain, as described in steps 5051-5054.
[0124] The embodiment of the application can accurately lock the specific target scene category of the autonomous vehicle by constructing a dynamic constraint domain, screening constraint adaptive interaction, forming a time sequence interaction event chain, extracting a static triggering feature, and performing comprehensive matching. It fully combines the space-time association and interaction relationship between the static element and the dynamic target, excludes subcategories that do not meet the scene features, makes the determined target scene category more suitable for the actual driving environment, provides accurate scene basis for the decision of the autonomous driving system, and helps to improve the response capability and safety of the autonomous vehicle in complex scenes.
[0125] In an embodiment, steps 5051-5054 are described as follows:
[0126] Step 5051, performing association analysis based on the environmental triggering signal of the static triggering feature and the behavior trajectory sequence of the dynamic target in the time sequence interaction event chain, and constructing a mapping relationship matrix; the mapping relationship matrix takes the dynamic target behavior mode as the row and the static triggering feature as the column; the mapping relationship matrix records the unique mapping of the static element trigger source and the dynamic target behavior corresponding to each interaction event.
[0127] Optionally, the scene classification device extracts the environmental trigger signals in the static trigger features (such as the color change of traffic lights, the position identification of zebra crossings, etc.) and the behavior trajectory sequences of dynamic targets in the time sequence interaction event chain (such as the path of pedestrians crossing the street, the acceleration / deceleration trajectory of vehicles, etc.), and performs correlation analysis on both, to clearly determine the corresponding relationship between the static element trigger source (such as traffic lights, zebra crossings) and the dynamic target behavior (such as pedestrian crossing, vehicle deceleration) in each interaction event. A mapping relationship matrix is constructed based on this, the rows of the matrix represent the behavior patterns of dynamic targets (such as “pedestrian crossing start” “vehicle deceleration avoidance” etc.), and the columns represent the static trigger features (such as “traffic light green light” “zebra crossing exists” etc.), and the elements in the matrix record the unique mapping of the static element trigger source and the dynamic target behavior corresponding to each interaction event, usually represented by 1 for existing mapping and 0 for non-existing mapping.
[0128] In an embodiment, the environmental trigger signals of the static trigger features include “traffic light green light” and “zebra crossing exists”, and the behavior patterns corresponding to the behavior trajectory sequences of dynamic targets are “pedestrian A crossing start” and “vehicle B deceleration avoidance”. Correlation analysis finds that “pedestrian A crossing start” is triggered by “traffic light green light”, and “vehicle B deceleration avoidance” is triggered by “zebra crossing exists”. The constructed mapping relationship matrix is as follows:
[0129] Dynamic target behavior patterns Traffic light green Zebra crossing present Pedestrian A starts crossing 1 0 Vehicle B slows down to avoid 0 1
[0130] The “1” in the matrix indicates that the corresponding behavior pattern and static trigger feature have a unique mapping, and the “0” indicates no mapping.
[0131] In step 5052, the frequency distribution of each static element trigger interaction event is counted based on the mapping relationship matrix, and the scene dynamic entropy is determined in combination with the type diversity of the interaction relationship between dynamic targets; the scene dynamic entropy represents the dynamic complexity of the scene under each subcategory.
[0132] Optionally, the scene classification device counts the frequency distribution of the interaction events triggered by each static element trigger source (i.e. static trigger feature) according to the mapping relationship matrix, for example, how many times the pedestrian crossing event is triggered by “traffic light green light”, how many times the vehicle deceleration event is triggered by “zebra crossing exists”, etc. At the same time, the type diversity of the interaction relationship between dynamic targets is analyzed, such as whether there are cross-avoidance, following driving, parallel driving and other interaction types. In combination with the frequency distribution and type diversity, the scene dynamic entropy is calculated through a preset dynamic entropy function, and the higher the entropy value, the higher the dynamic complexity of the scene under the subcategory (i.e. the more dispersed the distribution of static element trigger events and the more diverse the interaction types).
[0133] Specifically, the preset dynamic entropy function is:
[0134]
[0135] wherein, p i represents the proportion of the interaction event frequency of the ith static trigger feature; f represents the trigger frequency of the ith feature; n represents the number of static trigger features; m represents the number of types of interaction relationship between dynamic targets; and a represents a diversity weight coefficient (which can be 0.5). In the function, the first term is the frequency distribution entropy, the second term is the type diversity entropy, and the sum of the two is the scene dynamic entropy H.
[0136] In an embodiment, in the "urban intersection with pedestrian and vehicle cross avoidance" subcategory, the mapping relationship matrix shows that the "traffic signal green light" triggers 2 pedestrian crossing events, and the "zebra crossing exists" triggers 3 vehicle deceleration events, and the frequency distribution is [2, 3]. The interaction relationship type between dynamic targets includes cross avoidance (1 type). The calculation of the scene dynamic entropy needs to consider the uniformity of the frequency distribution and the type diversity. In this subcategory, the frequency distribution is relatively concentrated (more vehicle deceleration events), and the interaction type is single, so the scene dynamic entropy is low, and the calculated value is about 0.97 (the value range can be set to [0, 5], and the larger the value, the higher the complexity). In the "urban intersection with only vehicle traffic" subcategory, if the static trigger features are "traffic signal red light" and "lane line exists", they trigger 4 vehicle parking events and 5 vehicle lane keeping events respectively, and the interaction type is only following driving between vehicles (1 type), and the scene dynamic entropy is about 0.99.
[0137] In step 5053, the subcategories that meet the current scene dynamic complexity are screened according to the scene dynamic entropy, to obtain candidate scene subcategories.
[0138] Optionally, the scene classification device analyzes the dynamic complexity of the current actual scene (which can be estimated by the number and type of interaction events collected in real time), and compares it with the scene dynamic entropy of each subcategory. The subcategories whose scene dynamic entropy matches the dynamic complexity of the current scene, i.e., the difference between their entropy values is within a preset threshold range (such as ±0.3), are determined as candidate scene subcategories.
[0139] In an embodiment, the dynamic complexity of the current scene is estimated to correspond to a dynamic entropy of about 1.25. The scene dynamic entropy of the "urban intersection with pedestrian and vehicle cross avoidance" subcategory is 1.2, with a difference of 0.05 (within ±0.3); the scene dynamic entropy of the "urban intersection with only vehicle traffic" subcategory is 1.1, with a difference of 0.15 (within the range); and the scene dynamic entropy of another subcategory "urban intersection without signal light control" is 2.5, with a difference of 1.25 (exceeding the range). Therefore, the candidate scene subcategories screened are "urban intersection with pedestrian and vehicle cross avoidance" and "urban intersection with only vehicle traffic".
[0140] In step 5054, for each sub-category in the candidate scene sub-categories, the static trigger features are matched with the features of the static elements triggering the dynamic target behavior in the current scene, and the sub-category with the largest number of matched features is determined as the target scene category.
[0141] Optionally, the scene classification device matches the static trigger features of each sub-category in the candidate scene sub-categories (for example, the "urban intersection with pedestrian and vehicle cross-avoidance" can be subdivided into "cross-avoidance intersection with traffic light" and "cross-avoidance intersection without traffic light") with the features of the static elements triggering the dynamic target behavior in the current scene (for example, whether the current traffic signal is working, whether the zebra crossing is clear, etc.), and counts the number of matches between each sub-category and the current features. The sub-category with the largest number of matched features is determined as the target scene category.
[0142] In an embodiment, the sub-categories of the candidate scene sub-category "urban intersection with pedestrian and vehicle cross-avoidance" include "cross-avoidance intersection with traffic light" and "cross-avoidance intersection without traffic light". In the current scene, the features of the static elements triggering the dynamic target behavior are "there is a traffic signal and the light color is normal" and "the zebra crossing is clear". The static trigger features of "cross-avoidance intersection with traffic light" include "there is a traffic signal and it is working" and "there is a zebra crossing", and the number of matches with the current features is 2; the static trigger features of "cross-avoidance intersection without traffic light" are "there is no traffic signal" and "there is a zebra crossing", and the number of matches with the current features is 1. Therefore, the target scene category is determined as "cross-avoidance intersection with traffic light".
[0143] The embodiment of the present application further refines the accuracy of scene classification by constructing a mapping relationship matrix to determine the association between static triggers and dynamic behaviors, calculating the scene dynamic entropy to quantify the scene complexity, screening candidate sub-categories and matching sub-categories. Not only the mapping relationship between static and dynamic elements is considered, but also the dynamic complexity of the scene is combined, so that the determination of the target scene category is more in line with the subtle features of the actual scene, providing a more accurate environment description for the automatic driving system, which helps to improve the decision accuracy and safety of the vehicle in complex traffic scenes.
[0144] Further, the scene classification device for automatic driving provided by the present application is described below, and the scene classification device for automatic driving described below can be correspondingly referred to the scene classification method for automatic driving described above.
[0145] Optionally, referring to Figure 2 , Figure 2 is a structural schematic diagram of the scene classification device for automatic driving provided by the present application, and the scene classification device for automatic driving comprises:
[0146] The acquisition module 210 is configured to acquire static element data and dynamic target data in a current scene of an autonomous vehicle; the static element represents an object whose position does not change in the scene; the dynamic target represents an object whose position or state changes in the scene;
[0147] The initial scene determination module 220 is configured to extract and combine features of the static element data based on types of the static elements, to obtain combined features of the static elements, and to match the combined features of the static elements with a preset static scene type library to determine an initial scene category;
[0148] The graph construction module 230 is configured to generate motion trajectory data of each dynamic target based on position coordinates of each dynamic target at different time instants recorded in the dynamic target data, and to analyze the static element data and the motion trajectory data to construct a time sequence correlation graph of each static element and each dynamic target in the current scene;
[0149] The interaction relationship determination module 240 is configured to determine interaction relationships between the dynamic targets based on relative motion states between the dynamic targets identified based on the dynamic target data, in combination with motion speeds and directions of each dynamic target;
[0150] The scene category determination module 250 is configured to determine a target scene category of the autonomous vehicle based on the time sequence correlation graph and the interaction relationships for each subcategory in the initial scene category.
[0151] The embodiment of the present application can construct a time sequence correlation graph by combining the motion trajectory data generated based on the dynamic target data with the static element data, and determine interaction relationships between the dynamic targets based on relative motion states, speeds and directions, so as to deeply mine correlations between the dynamic targets and between the dynamic targets and the static environment, effectively solve the problem of insufficient analysis of dynamic element interactions in the prior art, in addition, the initial scene category can be refined and judged based on the time sequence correlation graph and the interaction relationships, so as to more comprehensively and accurately capture dynamic change characteristics of the scene, compared with the prior art which only relies on static features, the target scene category can be determined more timely and accurately, so as to improve the accuracy and real-time performance of scene classification, meet the high-precision requirement of scene judgment of autonomous driving, and more accurately identify potential risks in the scene, reduce decision-making errors caused by misjudgment of scene classification, and ensure the safety of autonomous driving.
[0152] Please refer to Figure 3 , Figure 3 An embodiment of an electronic device provided by the embodiment of the present application is shown in the figure. Figure 3As shown in the figure, the embodiment of the application provides an electronic device 300, which comprises a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and capable of running on the processor 320, and the processor 320 implements the following steps when executing the computer program 311:
[0153] Obtaining static element data and dynamic target data in a current scene of an autonomous vehicle; the static element represents an object whose position does not change in the scene; the dynamic target represents an object whose position or state changes in the scene;
[0154] Extracting and combining features of the static element data based on types of the static elements, obtaining static element combined features, and matching the static element combined features with a preset static scene type library to determine an initial scene category;
[0155] Generating motion trajectory data of each dynamic target based on position coordinates of each dynamic target at different time instants recorded in the dynamic target data, and analyzing the static element data and the motion trajectory data to construct a time sequence correlation graph of each static element and each dynamic target in the current scene;
[0156] Determining an interaction relationship between the dynamic targets based on relative motion states between the dynamic targets identified based on the dynamic target data, and combining a motion speed and a motion direction of each dynamic target;
[0157] For each subcategory in the initial scene category, determining a target scene category of the autonomous vehicle based on the time sequence correlation graph and the interaction relationship.
[0158] Please refer to Figure 4 , Figure 4 The embodiment of the computer readable storage medium provided by the embodiment of the application is shown in the figure. Figure 4 As shown in the figure, the embodiment provides a computer readable storage medium 400, which stores a computer program 311, and the computer program 311 is executed by a processor to implement the following steps:
[0159] Obtaining static element data and dynamic target data in a current scene of an autonomous vehicle; the static element represents an object whose position does not change in the scene; the dynamic target represents an object whose position or state changes in the scene;
[0160] Extracting and combining features of the static element data based on types of the static elements, obtaining static element combined features, and matching the static element combined features with a preset static scene type library to determine an initial scene category;
[0161] Based on the position coordinates of each dynamic target recorded in the dynamic target data at different time, motion trajectory data of each dynamic target is generated, and based on the static element data and the motion trajectory data, a time sequence correlation atlas of each static element and each dynamic target in the current scene is constructed;
[0162] Based on the relative motion state between the dynamic targets identified by the dynamic target data, in combination with the motion speed and motion direction of each dynamic target, an interaction relationship between the dynamic targets is determined.
[0163] For each subcategory in the initial scene category, based on the time sequence correlation atlas and the interaction relationship, a target scene category of the autonomous vehicle is determined.
[0164] On the other hand, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transient computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the scene classification method for autonomous driving provided by each method, the method comprising:
[0165] Obtaining static element data and dynamic target data in a current scene of an autonomous vehicle; the static element represents an object in the scene whose position does not change; the dynamic target represents an object in the scene whose position or state changes;
[0166] Based on the type of the static element, the static element data is extracted and combined to obtain a static element combined feature, and the static element combined feature is matched with a preset static scene type library to determine an initial scene category;
[0167] Based on the position coordinates of each dynamic target recorded in the dynamic target data at different time, motion trajectory data of each dynamic target is generated, and based on the static element data and the motion trajectory data, a time sequence correlation atlas of each static element and each dynamic target in the current scene is constructed;
[0168] Based on the relative motion state between the dynamic targets identified by the dynamic target data, in combination with the motion speed and motion direction of each dynamic target, an interaction relationship between the dynamic targets is determined.
[0169] For each subcategory in the initial scene category, based on the time sequence correlation atlas and the interaction relationship, a target scene category of the autonomous vehicle is determined.
[0170] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods of the embodiments or some parts of the embodiments.
[0172] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A scene classification method for automatic driving, characterized by, The method comprises the following steps: acquiring static element data and dynamic target data in a current scene of an autonomous vehicle; a static element represents an object whose position does not change in a scene; a dynamic target represents an object whose position or state changes in a scene; extracting and combining features of the static element data based on the type of the static element, obtaining static element combined features, and matching the static element combined features with a preset static scene type library to determine an initial scene category; generating motion trajectory data of each dynamic target based on position coordinates of each dynamic target at different time instants recorded in the dynamic target data, and analyzing the static element data and the motion trajectory data to construct a time sequence correlation graph of each static element and each dynamic target in the current scene; determining an interaction relationship between dynamic targets based on relative motion states between dynamic targets identified from the dynamic target data, and combining motion speed and motion direction of each dynamic target; determining a target scene category of the autonomous vehicle based on the time sequence correlation graph and the interaction relationship for each subcategory in the initial scene category; the step of determining the target scene category based on the time sequence correlation graph and the interaction relationship for each subcategory in the initial scene category comprises the following steps: constructing a dynamic constraint domain based on the time sequence correlation graph and the interaction relationship for each subcategory in the initial scene category; the dynamic constraint domain limits the motion range boundary of the dynamic target and the correlation dimension of the static element under each subcategory; determining a constraint adaptive interaction completely falling within the dynamic constraint domain based on the dynamic constraint domain and the interaction relationship between dynamic targets, and combining the matching degree of each interaction relationship within the dynamic constraint domain; serially connecting the constraint adaptive interaction and the time dimension of the time sequence correlation graph with the interaction relationship between dynamic targets in chronological order to form a time sequence interaction event chain under each subcategory; each interaction event in the time sequence interaction event chain comprises start and end time instants, a specific position where the interaction occurs, and a behavior result; extracting a static trigger feature of a trigger correlation between the dynamic target and the static element when each interaction event occurs for the time sequence interaction event chain; determining the target scene category based on the static trigger feature and the time sequence interaction event chain.
2. The scene classification method for automatic driving according to claim 1, characterized in that, the step of determining the target scene category based on the static trigger feature and the time sequence interaction event chain comprises the following steps: performing correlation analysis on an environmental trigger signal of the static trigger feature and a behavior trajectory sequence of the dynamic target in the time sequence interaction event chain to construct a mapping relationship matrix; the mapping relationship matrix takes a dynamic target behavior pattern as a row and a static trigger feature as a column; the mapping relationship matrix records a unique mapping between a static element trigger source and a dynamic target behavior corresponding to each interaction event; determining a scene dynamic entropy based on a frequency distribution of each static element trigger interaction event and combining the type diversity of the interaction relationship between dynamic targets; the scene dynamic entropy represents the dynamic complexity of the scene under each subcategory. screening a sub-category conforming to a current scene dynamic complexity according to the scene dynamic entropy, to obtain a candidate scene sub-category; for each sub-category in the candidate scene sub-category, matching the static trigger feature with a feature of a static element triggering a dynamic target behavior in a current scene, and determining a sub-category with the most matching features as the target scene category.
3. The scene classification method for automatic driving according to claim 1, characterized in that, the analysis based on the static element data and the motion trajectory data to construct a time sequence correlation graph of each static element and each dynamic target in the current scene, comprising: determining a spatial attribution relationship between each dynamic target and a static element at each time based on the position coordinates of the static element in the static element data and the motion trajectory data of the dynamic target, and constructing an initial correlation pair based on the spatial attribution relationship; extracting an instantaneous feature in the motion trajectory data of each dynamic target in each correlation pair and an interaction feature of the original attribute of the static element based on the initial correlation pair, to obtain an interaction feature value; statistically analyzing the motion feature coupling degree of different dynamic targets associated with the same static element based on the interaction feature value; analyzing the transfer strategy of the dynamic target between different static elements based on the initial correlation pair, to construct a time sequence constraint relationship of the static element to the dynamic target; constructing the time sequence correlation graph based on the time sequence constraint relationship, the motion feature coupling degree, and the interaction feature value.
4. The scene classification method for automatic driving according to claim 3, characterized in that, the analysis based on the time sequence constraint relationship, the motion feature coupling degree, and the interaction feature value to construct the time sequence correlation graph, comprising: performing constraint analysis based on the time sequence constraint relationship and the interaction feature value to obtain a time sequence dependent feature of the dynamic target in the static element sequence; the time sequence dependent feature reflects the response of the dynamic target behavior to the change of the static element; the static element sequence refers to a continuous sequence formed by the static elements associated with the dynamic target at different times in chronological order; performing constraint screening based on the motion feature coupling degree and the time sequence dependent feature to construct a multi-element interaction constraint network containing static elements and dynamic targets; combining the multi-element interaction constraint network and the motion feature coupling degree to obtain an association strength value of an edge in the graph; integrating the association relationship at all times based on the association strength value and the initial correlation pair to generate the time sequence correlation graph containing static elements, dynamic targets, and their time sequence correlation.
5. The scene classification method for automatic driving according to claim 1, characterized in that, the matching of the static element combination feature with the preset static scene type library to determine an initial scene category, comprising: structurally disassembling the static element combination feature to obtain multiple independent dimension features; each independent dimension corresponds to the original attribute of a type of static element; extracting the feature dimension information of the static elements contained in each scene type in the static scene type library to construct a scene type and dimension index table; statistically analyzing the matching degree of each dimension feature and the dimension index of each scene type based on the scene type and dimension index table to construct a dimension matching degree matrix; each element in the dimension matching degree matrix reflects the degree of fit between a certain feature dimension and a certain scene type in the dimension layer; Screening a scene type satisfying a preset dimension matching threshold based on the dimension matching degree matrix to obtain a candidate scene type; Matching and sorting each scene type in the candidate scene type to determine the initial scene category.
6. The scene classification method for automatic driving according to claim 5, characterized in that, The matching and sorting each scene type in the candidate scene type to determine the initial scene category comprises: For each scene type in the candidate scene type, extracting the feature details of each candidate scene type from the static scene type library to obtain scene type detail features; Comparing the dimension features of each dimension with the scene type detail features dimension by dimension to obtain the feature detail matching degree of each candidate scene type in each independent dimension; Integrating the feature detail matching degrees of each candidate scene type to obtain the feature detail total matching degree of each candidate scene type; Prioritizing each scene type in the candidate scene type based on the feature detail total matching degree from high to low, and determining the scene type at the top of the ranking as the initial scene category.
7. A scene classification apparatus for autonomous driving, characterized by, The scene classification method for automatic driving according to any one of claims 1 to 6; The scene classification device for automatic driving comprises: An acquisition module configured to acquire static element data and dynamic target data in a current scene of an automatic driving vehicle; a static element represents an object in the scene whose position does not change; a dynamic target represents an object in the scene whose position or state changes; An initial scene determination module configured to extract and combine features of the static element data based on the type of the static element to obtain static element combined features, and match the static element combined features with a preset static scene type library to determine an initial scene category; A graph construction module configured to generate motion trajectory data of each dynamic target based on position coordinates of each dynamic target at different time instants recorded in the dynamic target data, and analyze the static element data and the motion trajectory data to construct a time sequence association graph of each static element and each dynamic target in the current scene; An interaction relationship determination module configured to determine an interaction relationship between dynamic targets based on relative motion states between the dynamic targets identified based on the dynamic target data, in combination with a motion speed and a motion direction of each dynamic target; A scene category determination module configured to determine a target scene category of the automatic driving vehicle based on the time sequence association graph and the interaction relationship for each subcategory in the initial scene category. The determining a target scene category of the automatic driving vehicle based on the time sequence association graph and the interaction relationship for each subcategory in the initial scene category comprises: For each subcategory in the initial scene category, constructing a dynamic constraint domain based on a space-time association relationship between static elements and dynamic targets in the time sequence association graph; the dynamic constraint domain limits the motion range boundary of the dynamic target and the associated dimension of the static element under each subcategory. determine constraint adaptation interactions completely falling in the dynamic constraint field based on the interaction relationship between the dynamic constraint field and the dynamic target, and the matching degree of each interaction relationship in the dynamic constraint field; concatenate the interaction relationship between the dynamic targets in the time dimension of the constraint adaptation interaction and the time sequence interaction graph in chronological order to form a time sequence interaction event chain under each subcategory; each interaction event in the time sequence interaction event chain contains start and end time, specific location and behavior result of interaction occurrence; extract the trigger association between the dynamic target and the static element when each interaction event occurs for the time sequence interaction event chain to obtain a static trigger feature; determine the target scene category based on the static trigger feature and the time sequence interaction event chain.
8. An electronic device, comprising: a memory for storing a computer software program; a processor for reading and executing the computer software program, wherein the processor, when executing the computer software program, implements the scene classification method for automatic driving according to any one of claims 1 to 6.
9. A non-transitory computer readable storage medium having stored therein a computer software program, characterized in that, the computer software program, when executed by the processor, implements the scene classification method for automatic driving according to any one of claims 1 to 6.