A driving scene understanding method based on spatial evidence constraint
By constructing spatial evidence units and reliability verification mechanisms, the problem of lack of spatial basis in the understanding of driving scenarios in existing technologies is solved, and reliable output and improved interpretability of results are achieved under complex conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-05-18
- Publication Date
- 2026-06-12
Smart Images

Figure CN122200604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving environment perception and scene understanding technology, and in particular to a driving scene understanding method based on spatial evidence constraints. Background Technology
[0002] With the development of autonomous driving technology, driving scene understanding has become a crucial processing step after environmental perception in autonomous driving systems. Existing scene understanding technologies typically first acquire environmental data around the vehicle using sensors such as cameras, LiDAR, and millimeter-wave radar. This environmental data is then processed through target detection, tracking, semantic segmentation, rasterization, vectorization, or occupancy representation. Finally, based on rules, scene models, neural network models, or multimodal models, the technologies output scene categories, risk descriptions, semantic conclusions, or question-and-answer results. Essentially, based on the perception results, they provide semantic interpretation and scene judgment of the current road environment. Existing technologies mainly include the following categories: One type of method focuses on scene recognition or classification. These methods typically label the current environment based on driving data, traffic participant behavior characteristics, trajectory information, or pre-defined scene models. For example, they might identify the environment as a following scenario, a meeting scenario, an intersection crossing scenario, or a lane-changing scenario. While suitable for testing and evaluation, scene statistics, and rule triggering, their outputs are mostly predefined categories or scene labels, making it difficult to express fine-grained information such as the complex local spatial relationships between road participants, the state of occluded areas, the boundaries of passable spaces, and potential risk areas.
[0003] Another type of approach focuses on structuring the perception results to serve subsequent prediction, planning, or control modules. For example, it converts perception results such as targets, lane lines, and road boundaries into vectorized features, or uses grids, semantic grids, or occupancy grids to spatially represent the environment. This type of approach improves information organization efficiency and provides a unified input for subsequent modules; however, its primary function remains as an internal intermediate representation. It typically does not directly address the interpretable output of scene understanding results, nor does it establish an explicit correspondence between the understanding results and specific spatial data.
[0004] With the development of multimodal modeling technology, a new approach has emerged that incorporates bird's-eye view features, occupancy representations, or other spatial representations into scene understanding models to enhance question answering and semantic reasoning capabilities. This type of method has certain advantages in open-ended descriptions, scene question answering, and complex semantic reasoning, and can improve the expressive power of driving scene understanding to some extent. However, this type of method typically focuses more on improving answering ability or task performance. Spatial representations often participate in reasoning in the form of implicit features, and the final output of the model still mainly represents textual conclusions or answers, lacking explicit correspondence between specific spatial regions and specific evidence objects.
[0005] In complex driving scenarios, especially under conditions of obstruction, severe weather, weak long-distance observation, cross-modal conflict, or lack of local information, existing technologies still have the following problems: First, existing scene recognition or scene question answering methods usually directly output scene conclusions, risk descriptions or text answers, but do not simultaneously provide specific spatial evidence to support the conclusions. This makes it difficult for the output results to be pointed back to the corresponding spatial region, target object or sensing unit, which is not conducive to result verification, anomaly localization and project acceptance. Secondly, while existing occupancy representations, grid representations, or vectorized features can describe the spatial state of the environment, they are mostly used as internal inputs for subsequent modules and have not yet been further organized into evidentiary units that can directly support the output of scene understanding. In other words, existing technologies can generally represent spatial states, but cannot yet constrain scene understanding results with spatial evidence; Furthermore, when perceptual input degrades or conflicts occur, existing model-based scene understanding methods may still output conclusions lacking sufficient perceptual support. This is because existing technologies generally lack evidentiary constraint mechanisms for understanding results, lack joint verification of the sufficiency, consistency, and reliability of evidence, and lack processing mechanisms to perform downgraded output or refuse to answer when evidence is insufficient.
[0006] Furthermore, existing technologies lack sufficient definition of the credibility boundaries for scene understanding results. Even if the system already possesses certain uncertainty estimates or occupancy probability information, this information is typically not explicitly incorporated into the scene understanding output process. This makes it difficult for the system to distinguish between scenarios that can provide definitive conclusions and those that can only offer risk warnings or require rejection. Therefore, how to establish an explicit correspondence between driving scene understanding conclusions and referentially verifiable spatial evidence, based on existing multimodal perception and spatial representation results, and how to execute degraded output or rejection output when evidence is insufficient, perception degrades, or cross-modal conflicts, is a pressing technical problem that needs to be solved in this field. Summary of the Invention
[0007] The purpose of this invention is to provide a driving scene understanding method based on spatial evidence constraints, in order to solve the technical problems in the prior art where the driving scene understanding results lack an explicit correspondence with the perceived spatial basis, and where conclusions lacking evidence support are easily output under conditions of occlusion, perception degradation, cross-modal conflict or information loss, and where it is difficult to constrain the reliability of the output results.
[0008] To achieve the above objectives, the present invention provides the following technical solution: A driving scenario understanding method based on spatial evidence constraints includes the following steps: S1. Acquire multimodal perception data of the environment surrounding the target vehicle, and perform time synchronization, spatial registration and coordinate alignment on the multimodal perception data to obtain perception input data under a unified vehicle coordinate system. S2. Construct an occupancy representation of the target scene based on the perceived input data, and map the perceived input data to preset spatial units to obtain the occupancy status corresponding to multiple spatial units; S3. Based on the differences in the representation of the same spatial unit in at least two observation sources, determine the consistency residual of the corresponding spatial unit; S4. Determine the uncertainty value of the spatial unit; the uncertainty value is used to characterize the reliability of the state of the corresponding spatial unit; S5. Construct spatial evidence units based on the spatial location, occupancy status, consistency residual and uncertainty value of each spatial unit, establish the mapping relationship between spatial evidence units and original sensing data, select target spatial evidence units from multiple spatial evidence units, and encode the target spatial evidence units to form a spatial evidence sequence. S6. Input the spatial evidence sequence and scene understanding request into the scene understanding model to generate candidate structured scene understanding results; the candidate structured scene understanding results include scene understanding conclusions, corresponding spatial evidence identifiers, and confidence values; S7. Based on the consistency residuals, uncertainty values, and consistency between conclusions and evidence of the spatial evidence corresponding to the spatial evidence identifier, verify the reliability of the candidate structured scene understanding results and determine the output mode to obtain the final structured scene understanding results. The output modes include normal output, downgraded output, and rejection output.
[0009] Furthermore, the occupancy state in S2 is used to characterize the corresponding spatial unit as being in one or more of the states of being occupied, passable, and unknown. The information represented by the occupancy state includes one or more of the following: object semantic information, drivable attribute information, and motion attribute information.
[0010] Furthermore, the preset spatial unit forms in S2 include: region-level spatial units based on bird's-eye view grids, voxel-level spatial units based on three-dimensional voxels, object-level spatial units based on object instances, and relation-level spatial units based on object relationships; spatial evidence units adopt region-level, voxel-level, object-level, and relation-level evidence forms respectively.
[0011] Furthermore, the consistency residual of the corresponding spatial unit is determined in S3 as follows: For the i Each spatial unit, consistency residual Differences in occupied states Semantic differences and geometric difference terms Weighted average yields: ; In the formula, , , These are the weighting coefficients for the corresponding items.
[0012] Furthermore, the uncertainty value for determining the spatial unit in S4 is specifically as follows: For the i Each spatial unit, uncertainty value According to the occupancy probability entropy Variance of results from multiple inferences and sensor mass degradation term Jointly determined: ; In the formula, , , These are the weighting coefficients for the corresponding items.
[0013] Furthermore, the mapping relationship between spatial evidence units and raw sensing data in S5 includes: mapping spatial evidence units to corresponding image regions in the raw image; mapping spatial evidence units to point cloud index sets or local point cloud regions in the raw point cloud; and mapping spatial evidence units to target numbers, azimuth ranges, or range intervals of millimeter-wave radar.
[0014] Furthermore, the specific steps for selecting target spatial evidence units from multiple spatial evidence units in S5 are as follows: For the i Each spatial evidence unit is scored based on its relevance to the scene understanding request. Consistency residuals and uncertainty value Calculate the overall screening score : ; In the formula, , , The weighting coefficients for the corresponding items are determined based on the overall screening score. Multiple spatial evidence units are sorted, and the spatial evidence units with the highest scores are selected to form a spatial evidence sequence.
[0015] Furthermore, the reliability verification of the candidate structured scene understanding results in S7 is specifically as follows: For the candidate structured scene understanding results, its reliability score Represented as: ; In the formula, , , These are the weighting coefficients for the corresponding items. This represents the average consistency residual of the cited spatial evidence. This represents the average uncertainty value of the cited spatial evidence. The score indicates the consistency between the conclusion and the evidence.
[0016] Furthermore, a first reliability threshold is set in S7. Second reliability threshold ,and Greater than When reliability score Greater than or equal to When the reliability score is high, the output mode is set to normal output; when the reliability score is high, the output mode is set to normal output. Less than and greater than or equal to When the reliability score is high, the output mode is determined to be degraded output; when the reliability score is low, the output mode is determined to be degraded output. Less than When the output mode is determined to be a rejection output, the rejection output indicates that the current spatial evidence is insufficient to support a definitive conclusion.
[0017] Furthermore, the final structured scene understanding result includes scene understanding conclusions, corresponding spatial evidence identifiers, confidence values, and output patterns. The downgraded output includes at least one of the following: coarse-grained scene description, risk warning information, and uncertain area identifiers.
[0018] As can be seen from the above technical solution, compared with the prior art, the present invention provides a driving scene understanding method based on spatial evidence constraints, which has the following beneficial effects: 1. This invention constructs spatial evidence units based on spatial location, occupancy status, consistency residuals, and uncertainty values, and establishes a mapping relationship between spatial evidence units and original perceptual data. This enables scene understanding results to no longer be presented as a single textual conclusion, but to be able to point back to the perceptual basis corresponding to specific spatial regions, object instances, grid units, or voxel units. This is beneficial to improving the interpretability and verifiability of driving scene understanding results and solving the problems of difficulty in verifying understanding conclusions and locating evidence in the prior art. 2. This invention can output structured results that include scene understanding conclusions, corresponding spatial evidence identifiers, confidence values, and output modes, so that an explicit correspondence is formed between the scene understanding conclusions and their supporting evidence, and the system can distinguish between three states: normal output, degraded output, and rejection output. This establishes a clear reliability boundary for the scene understanding results and avoids the problem in the prior art of only outputting conclusions without supporting evidence and state descriptions. 3. This invention introduces a consistency residual, uncertainty value, and consistency verification mechanism between conclusions and evidence. In the case of occlusion, perceptual degradation, cross-modal conflict, or lack of local information, it can perform downgraded output or reject output for candidate results that lack sufficient evidence support, instead of continuing to output definitive conclusions. This helps to suppress erroneous understanding results that lack spatial evidence support in complex scenarios and improves the credibility and robustness of scenario understanding results. 4. This invention filters spatial evidence units according to the scene understanding request and organizes the filtered target spatial evidence units into a spatial evidence sequence for input into the scene understanding model. This can reduce the interference of irrelevant perception information on the scene understanding process, highlight key spatial areas and risk areas related to the current understanding task, and help improve the pertinence of driving scene understanding processing and information utilization efficiency. Attached Figure Description
[0019] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the embodiments of the present invention are briefly described below. The following drawings are all schematic diagrams and are only used to illustrate the technical concept and main technical features of the present invention, and do not constitute a limitation on the scope of protection of the present invention. The same reference numerals in the drawings represent the same or corresponding technical features.
[0020] Figure 1 This is a flowchart illustrating the overall process of the driving scenario understanding method based on spatial evidence constraints according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the spatial evidence construction and mapping relationship in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the generation and reliability verification of scenario understanding results in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] This invention discloses a driving scenario understanding method based on spatial evidence constraints, such as... Figure 1 As shown, it includes the following steps: The method disclosed in this invention can be deployed on the vehicle's onboard computing platform, roadside edge computing device, or server to receive multimodal perception data of the vehicle's surrounding environment and output driving scene understanding results; it only processes the driving scene understanding process after perception and does not involve trajectory planning, vehicle control, or execution decision output.
[0023] S1. Acquire multimodal perception data of the environment surrounding the target vehicle, and perform time synchronization, spatial registration and coordinate alignment on the multimodal perception data to obtain perception input data under a unified vehicle coordinate system; provide a unified input basis for subsequent occupation representation construction, spatial evidence generation and scene understanding, and reduce the impact of inconsistent sampling times and coordinate systems of different sensors on the results; S2. Based on the perceived input data, construct the occupancy representation of the target scene, map the perceived input data to preset spatial units, and obtain the occupancy state corresponding to multiple spatial units; S2 converts the heterogeneous perception results into a unified spatial basic representation, providing a spatial carrier for subsequent consistency analysis and evidence organization; S3. Based on the differences in the representation of the same spatial unit in at least two observation sources, determine the consistency residual of the corresponding spatial unit; S3 can explicitly identify spatial regions with perceptual conflicts or potential risks of misunderstanding. S4. Determine the uncertainty value of the spatial unit; the uncertainty value is used to characterize the reliability of the state of the corresponding spatial unit; S4 can distinguish between spatial regions with sufficient and reliable information and spatial regions with insufficient information or low reliability. S5, such as Figure 2 As shown, spatial evidence units are constructed based on the spatial location, occupancy status, consistency residual, and uncertainty value of each spatial unit. A mapping relationship between spatial evidence units and original sensing data is established. Target spatial evidence units are selected from multiple spatial evidence units and encoded to form a spatial evidence sequence. S5 can organize the underlying sensing results into explicit evidence carriers that are referential, verifiable, and usable for scene understanding. S6. Input the spatial evidence sequence and scene understanding request into the scene understanding model to generate candidate structured scene understanding results. The candidate structured scene understanding results include scene understanding conclusions, corresponding spatial evidence identifiers, and confidence values. S6 enables the scene understanding conclusions to establish an explicit correspondence with specific spatial evidence, rather than just outputting unverifiable text conclusions. S7, such as Figure 3 As shown, based on the consistency residuals, uncertainty values, and consistency between conclusions and evidence corresponding to spatial evidence identifiers, the reliability of candidate structured scene understanding results is verified, and output modes are determined to obtain the final structured scene understanding results. Output modes include normal output, downgraded output, and rejection output. When evidence is insufficient, conflicts are obvious, or reliability is low, conclusions lacking spatial basis are suppressed, establishing clear reliability boundaries for the scene understanding results.
[0024] In this embodiment of the invention, the multimodal perception data includes at least two of image data, lidar point cloud data, and millimeter-wave radar data. Furthermore, it also includes one or more of vehicle speed, acceleration, heading angle, odometer information, and inertial measurement unit data. The observation sources in S3 include one or more of the following: different sensor modes, different viewpoints, different time segments, and reference observations obtained by backprojection or reconstruction based on occupancy characterization; the consistency residual is used to characterize the degree of occupancy conflict, geometric inconsistency, or semantic inconsistency of the corresponding spatial unit among different observation sources. Each spatial evidence unit includes at least a spatial identifier, spatial location, occupancy status, consistency residual, and uncertainty value; it may also include one or more of the following: semantic category, source modality identifier, and time identifier. In S6, scene understanding requests can be one or more of natural language questions, structured task instructions, or pre-defined scene summarization tasks. The scene understanding model can be a multimodal large language model or other semantic reasoning model, and serves only as a post-perception scene understanding service layer, not for outputting trajectory planning results, vehicle control variables, or driving execution strategies. The scene understanding model can be trained using training samples containing standard scene understanding conclusions, standard spatial evidence annotations, and output pattern labels. This allows the scene understanding model to learn to generate scene understanding conclusions corresponding to spatial evidence, and to learn to output downgraded or rejected results in scenarios with high residuals, high uncertainty, or insufficient information.
[0025] Furthermore, the occupancy state in S2 is used to characterize the corresponding spatial unit as being in one or more of the states of being occupied, passable, and unknown. The information represented by the occupancy state includes one or more of the following: object semantic information, drivable attribute information, and motion attribute information.
[0026] Furthermore, the preset spatial unit forms in S2 include: region-level spatial units based on bird's-eye view grids, voxel-level spatial units based on three-dimensional voxels, object-level spatial units based on object instances, and relation-level spatial units based on object relationships; spatial evidence units adopt region-level, voxel-level, object-level, and relation-level evidence forms respectively.
[0027] Furthermore, the consistency residual of the corresponding spatial unit is determined in S3 as follows: For the i Each spatial unit, consistency residual Differences in occupied states Semantic differences and geometric difference terms Weighted average yields: ; In the formula, , , These are the weighting coefficients for the corresponding items; This indicates the degree of difference in the occupancy status of the spatial unit from different observation sources. This indicates the degree of difference in the semantic category of the spatial unit from different observation sources. This indicates the degree of geometric deviation of the spatial unit after reprojection, reconstruction, or motion compensation. In this embodiment, all weights are greater than 0. Through the above method, cross-modal conflicts, semantic inconsistencies, and geometric deviations can be uniformly quantified into the degree of spatial risk.
[0028] For example, in one embodiment of the present invention, when image data indicates that a certain spatial region is passable, and the lidar point cloud shows a stable echo in that region, Increase; when the image semantic segmentation and point cloud semantic classification results are inconsistent. Increase; when significant spatial drift still exists between adjacent time segments after motion compensation. Increase.
[0029] Furthermore, the uncertainty value for determining the spatial unit in S4 is specifically as follows: For the i Each spatial unit, uncertainty value According to the occupancy probability entropy Variance of results from multiple inferences and sensor mass degradation term Jointly determined: ; In the formula, , , These are the weighting coefficients for the corresponding items. In this embodiment, all weights are greater than 0. Through the above method, a more reasonable uncertainty assessment can be given for obstructed areas, distant weak observation areas, and areas degraded by severe weather.
[0030] In this embodiment of the invention, the first i Each spatial evidence unit It can be represented as: ; In the formula, Number the spatial evidence. For spatial location, To occupy the state, For consistent residuals, For uncertain values, For semantic categories, As a source modality identifier, the above structure enables spatial evidence units to not only represent "what the region is", but also "whether there is a conflict in the region" and "whether the region is reliable".
[0031] Furthermore, the mapping relationship between spatial evidence units and raw sensing data in S5 includes: mapping spatial evidence units to corresponding image regions in the raw image; mapping spatial evidence units to point cloud index sets or local point cloud regions in the raw point cloud; and mapping spatial evidence units to target numbers, azimuth ranges, or range intervals of millimeter-wave radar. By establishing this mapping relationship, subsequent scene understanding results can be referenced back to the original sensing data.
[0032] In this embodiment of the invention, for image data, spatial evidence units can be mapped to corresponding image regions through a camera projection matrix; for point cloud data, a correspondence can be established between spatial evidence units and point cloud index sets or local point cloud regions; for millimeter-wave radar data, a correspondence can be established between spatial evidence units and target numbers, azimuth ranges, or distance intervals.
[0033] Furthermore, the specific steps for selecting target spatial evidence units from multiple spatial evidence units in S5 are as follows: For the i Each spatial evidence unit is scored based on its relevance to the scene understanding request. Consistency residuals and uncertainty value Calculate the overall screening score : ; In the formula, , , The weighting coefficients for the corresponding items are determined based on the overall screening score. Multiple spatial evidence units are sorted, and the highest-scoring units are selected to form a spatial evidence sequence. In this embodiment, all weights are greater than 0. This method highlights key spatial evidence that is relevant to the scene understanding request and has high risk warning significance, while reducing interference from irrelevant information.
[0034] Furthermore, the reliability verification of the candidate structured scene understanding results in S7 is specifically as follows: For the candidate structured scene understanding results, its reliability score Represented as: ; In the formula, , , These are the weighting coefficients for the corresponding items. This represents the average consistency residual of the cited spatial evidence. This represents the average uncertainty value of the cited spatial evidence. This represents the consistency score between the conclusion and the evidence. The above method allows for explicit reliability boundary control of the scene understanding results. The consistency score is determined based on the degree of matching between the object categories, spatial locations, regional attributes, or object relationships involved in the candidate structured scene understanding results and the semantic categories, spatial locations, occupancy states, or source modality identifiers in the cited spatial evidence units.
[0035] Furthermore, a first reliability threshold is set in S7. Second reliability threshold ,and Greater than When reliability score Greater than or equal to When the reliability score is high, the output mode is set to normal output; when the reliability score is high, the output mode is set to normal output. Less than and greater than or equal to When the reliability score is high, the output mode is determined to be degraded output; when the reliability score is low, the output mode is determined to be degraded output. Less than When the output mode is determined to be a rejection output, the rejection output indicates that the current spatial evidence is insufficient to support a definitive conclusion.
[0036] In this embodiment of the invention, the first reliability threshold and the second reliability threshold can be determined in any of the following ways: determined based on a fixed threshold, determined adaptively based on scene category, weather conditions, sensor status or road type, or determined jointly based on historical statistical results and online calibration results.
[0037] Furthermore, the final structured scene understanding result includes scene understanding conclusions, corresponding spatial evidence identifiers, confidence values, and output patterns. The downgraded output includes at least one of the following: coarse-grained scene description, risk warning information, and uncertain area identifiers.
[0038] In this embodiment of the invention, the scene understanding model first encodes the spatial evidence sequence, for each spatial evidence unit... Generate corresponding evidence vectors The evidence vectors include at least spatial location encoding, occupancy state embedding, consistency residual embedding, and uncertainty value embedding, and may further include semantic category embedding and source modality embedding. Then, multiple evidence vectors are arranged in a preset order to form a spatial evidence sequence, which is input into the scene understanding model along with the scene understanding request. The scene understanding model outputs candidate structured scene understanding results. For example, when the scene understanding request is "Is there a potential risk affecting the straight-ahead movement of the vehicle to the right?", the candidate structured scene understanding results can be represented as: There is an occlusion risk behind the bus to the right; it cannot be confirmed whether there is a dynamic target within the occluded area; the corresponding spatial evidence identifiers are E31, E32, E35, and E36; the confidence value is 0.63. In one implementation, the confidence value in the final structured scene understanding result can be expressed as a reliability score. Alternatively, a reliability score can be used. The confidence value obtained after calibrating the candidate confidence values.
[0039] The target vehicle is equipped with cameras, lidar, and millimeter-wave radar to collect multimodal perception data of the surrounding environment. The perception data is uniformly transformed into the vehicle coordinate system based on the intrinsic and extrinsic parameters of each sensor. For vehicles in motion, motion compensation can be performed based on the vehicle's state data to reduce the impact of inconsistent sampling times of different modalities on subsequent spatial representation.
[0040] In one embodiment of the invention, the target vehicle approaches an urban intersection, where a bus on its right front obstructs the pedestrian crossing entrance area. The camera can capture images of the bus itself and part of the road area, and the lidar can obtain point cloud echoes of the visible portion of the bus, but it lacks effective observation of the obstructed area behind the bus. The millimeter-wave radar in this area only obtains sparse target reflections.
[0041] In this scenario, the area where the bus itself is located is marked as occupied, and the area behind it that is obscured is marked as unknown. Significant inconsistencies are identified between the images, point clouds, and radar observations near the obscuration boundary. Based on the occupancy probability entropy and sensor quality attenuation terms, it is determined that the obscured area has high uncertainty. Spatial evidence units corresponding to the bus itself, the obscuration boundary, and the pedestrian crossing entrance are constructed, forming a target spatial evidence sequence.
[0042] When the scenario understanding request is "Is there a potential risk affecting the vehicle's straight-ahead movement to the right?", candidate structured scenario understanding results are generated based on the target space evidence sequence. If the average consistency residual and average uncertainty value of the cited spatial evidence are high, and the consistency between the conclusion and the evidence is insufficient to support a definitive judgment, the output mode is determined to be a downgraded output or a rejection output, instead of directly outputting a definitive conclusion of "no risk" or "pedestrians present".
[0043] For example, when the reliability is in the degraded output range, the final structured scenario understanding result can be represented as: Scenario understanding conclusion: There is an obstructed area behind the bus on the right front, which may pose a potential dynamic risk to the straight-ahead traffic. Corresponding spatial evidence identifiers: E31, E32, E35, E36; Confidence level: 0.58; Output mode: Degraded output.
[0044] When the obstructed area is further compounded by severe weather or weak observation from a distance, causing the reliability score to fall below the rejection threshold, the final structured scene understanding result can be expressed as: Scene understanding conclusion: The current spatial evidence is insufficient to determine whether there is a dynamic target in the area obscured to the right front; Corresponding spatial evidence identifiers: E31, E32, E35, E36; Confidence level: 0.32; Output mode: Rejected output.
[0045] The technical terms used in this embodiment have the following meanings: Occupied status is used to indicate whether an environmental space unit is occupied, accessible, or unknown. Consistency residuals refer to the amount of difference in the occupancy status, semantic category, or geometric location of the same spatial unit across different observation sources. Uncertainty value refers to a quantitative indicator used to characterize the reliability of the state of a corresponding spatial unit; Spatial evidence refers to evidence carriers that consist of spatial location, occupancy status, consistency residuals, uncertainty values, and optional semantic information, and can be traced back to the original perceived data.
[0046] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0047] Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A driving scenario understanding method based on spatial evidence constraints, characterized in that, Includes the following steps: S1. Acquire multimodal perception data of the environment surrounding the target vehicle, and perform time synchronization, spatial registration and coordinate alignment on the multimodal perception data to obtain perception input data under a unified vehicle coordinate system. S2. Construct an occupancy representation of the target scene based on the perceived input data, and map the perceived input data to preset spatial units to obtain the occupancy status corresponding to multiple spatial units; S3. Based on the differences in the representation of the same spatial unit in at least two observation sources, determine the consistency residual of the corresponding spatial unit; S4. Determine the uncertainty value of the spatial unit; the uncertainty value is used to characterize the reliability of the state of the corresponding spatial unit; S5. Construct spatial evidence units based on the spatial location, occupancy status, consistency residual and uncertainty value of each spatial unit, establish the mapping relationship between spatial evidence units and original sensing data, select target spatial evidence units from multiple spatial evidence units, and encode the target spatial evidence units to form a spatial evidence sequence. S6. Input the spatial evidence sequence and scene understanding request into the scene understanding model to generate candidate structured scene understanding results; The candidate structured scene understanding results include scene understanding conclusions, corresponding spatial evidence identifiers, and confidence values; S7. Based on the consistency residuals, uncertainty values, and consistency between conclusions and evidence of the spatial evidence corresponding to the spatial evidence identifier, verify the reliability of the candidate structured scene understanding results and determine the output mode to obtain the final structured scene understanding results. The output modes include normal output, downgraded output, and rejection output.
2. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, In S2, the occupancy state is used to represent one or more of the following states: occupied, passable, and unknown. The information represented by the occupancy state includes one or more of the following: object semantic information, drivable attribute information, and motion attribute information.
3. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The preset spatial unit forms in S2 include: regional spatial units based on bird's-eye view grids, voxel-level spatial units based on three-dimensional voxels, object-level spatial units based on object instances, and relation-level spatial units based on object relationships; spatial evidence units adopt corresponding regional, voxel-level, object-level, and relation-level evidence forms.
4. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The consistency residuals for the corresponding spatial units are determined in S3 as follows: For the i Each spatial unit, consistency residual Differences in occupied states Semantic differences and geometric difference terms Weighted average yields: ; In the formula, , , These are the weighting coefficients for the corresponding items.
5. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The uncertainty value for determining the spatial unit in S4 is specifically as follows: For the i Each spatial unit, uncertainty value According to the occupancy probability entropy Variance of results from multiple inferences and sensor mass degradation term Jointly determined: ; In the formula, , , These are the weighting coefficients for the corresponding items.
6. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The mapping relationship between spatial evidence units and raw sensing data in S5 includes: mapping spatial evidence units to corresponding image regions in the raw image; mapping spatial evidence units to point cloud index sets or local point cloud regions in the raw point cloud; and mapping spatial evidence units to target numbers, azimuth ranges, or range intervals of millimeter-wave radar.
7. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The specific steps for selecting a target spatial evidence unit from multiple spatial evidence units in S5 are as follows: For the i Each spatial evidence unit is scored based on its relevance to the scene understanding request. Consistency residuals and uncertainty value Calculate the overall screening score : ; In the formula, , , The weighting coefficients for the corresponding items are determined based on the overall screening score. Multiple spatial evidence units are sorted, and the spatial evidence units with the highest scores are selected to form a spatial evidence sequence.
8. The driving scenario understanding method based on spatial evidence constraints according to claim 1, characterized in that, The reliability verification of the candidate structured scene understanding results in S7 is specifically as follows: For the candidate structured scene understanding results, its reliability score Represented as: ; In the formula, , , These are the weighting coefficients for the corresponding items. This represents the average consistency residual of the cited spatial evidence. This represents the average uncertainty value of the cited spatial evidence. The score indicates the consistency between the conclusion and the evidence.
9. The driving scenario understanding method based on spatial evidence constraints according to claim 8, characterized in that, Set a first reliability threshold in S7 Second reliability threshold ,and Greater than When reliability score Greater than or equal to When the reliability score is high, the output mode is set to normal output; when the reliability score is high, the output mode is set to normal output. Less than and greater than or equal to When the reliability score is high, the output mode is determined to be degraded output; when the reliability score is low, the output mode is determined to be degraded output. Less than When the output mode is determined to be a rejection output, the rejection output indicates that the current spatial evidence is insufficient to support a definitive conclusion.
10. The driving scene understanding method based on spatial evidence constraints according to claim 9, characterized in that, The final structured scene understanding results include scene understanding conclusions, corresponding spatial evidence identifiers, confidence values, and output patterns. The downgraded output includes at least one of the following: coarse-grained scene description, risk warning information, and uncertain area identifiers.