A traffic scene understanding method, device, medium and product
Patent Information
- Application Number
- CN202511261352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-09-04
AI Technical Summary
[0005]本申请提供了一种交通场景理解方法、设备、介质及产品,以至少解决相关技术中缺乏对交通环境的深层次和全局性理解,模型无法有效泛化到新的交通环境,从而增加车辆行驶风险的问题
[0010] Beneficial Effects: This application addresses the current approach, which focuses solely on understanding specific target categories and lacks a comprehensive and in-depth understanding of traffic scenarios. Therefore, after extracting image features from the current road scene image, this application sequentially obtains the detection results and trajectory information of the target object through target detection and trajectory tracking. The detection results, trajectory information, and historical scene understanding results are then input into a pre-trained scene structuring model to obtain structured data of the current road scene, thereby achieving structured prediction of the road scene and significantly improving the understanding performance of subsequent scene understanding models. Furthermore, by incorporating historical scene understanding results into the current scene structure prediction, the continuity of time can compensate for the information loss in single-frame data, significantly improving the stability of understanding in complex scenarios and solving the problem of poor adaptability of traditional techniques to sudden scenarios. Further, this application also performs target behavior recognition based on the detection results, trajectory information, and structured data to obtain the behavior recognition results of the target object. Finally, by fusing multi-dimensional information such as image features, detection results, and behavior recognition results through a scene understanding model, the current scene understanding result is output, thereby gaining a more comprehensive and in-depth understanding of the traffic scene, achieving the technical effect of improving the accuracy of traffic scene understanding and ensuring vehicle driving safety.
Smart Images

Figure CN120953962B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method, device, medium, and product for understanding traffic scenes. Background Technology
[0002] With the rapid development and widespread application of autonomous driving technology, autonomous vehicles inevitably coexist with human-driven vehicles in complex mixed traffic scenarios. In this context, scene understanding and intent recognition are key technologies for detecting and tracking critical road targets, and are crucial for ensuring the safe operation of vehicles in complex traffic environments.
[0003] Currently, it is possible to detect and track road targets such as vehicles, pedestrians, and traffic signs, as well as predict the behavior and trajectories of traffic participants such as vehicles. However, models typically focus only on understanding specific target categories, such as pedestrians or vehicles, or specific behaviors, such as driving against traffic or changing lanes, lacking a global understanding of the traffic scenario. This makes it difficult to interpret their predictions, limiting the ability of autonomous driving models and computing systems to cope with complex environments. Due to the lack of a deep understanding of the traffic environment, models may not be able to effectively generalize to new traffic environments, thereby increasing the risk of vehicle driving.
[0004] In summary, how to gain a more comprehensive and in-depth understanding of traffic scenarios, thereby improving the accuracy of traffic scenario understanding and ensuring vehicle driving safety, is a problem that needs to be solved. Summary of the Invention
[0005] This application provides a traffic scene understanding method, device, medium, and product to at least address the problem in related technologies that lack a deep and global understanding of the traffic environment, and that models cannot be effectively generalized to new traffic environments, thereby increasing the risk of vehicle driving.
[0006] This application provides a method for understanding traffic scenarios, including: Feature extraction is performed on the current road scene image acquired by the sensor to obtain image features, and target detection is performed on the image features to obtain the detection result of the target object; Based on the detection results, the trajectory of the target object is tracked and identified to obtain the trajectory information of the target object; The detection results, trajectory information, and historical scene understanding results are input into a pre-trained scene structuring model to obtain structured data of the current road scene; the scene structuring model is a model used to perform structuring processing on road scene data; The behavior of the target object is identified based on the detection results, trajectory information, and structured data to obtain the behavior identification results of the target object; Image features, detection results, and behavior recognition results are input into a pre-trained scene understanding model to obtain scene understanding results for the current road scene.
[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the traffic scene understanding methods described above.
[0008] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the traffic scene understanding methods described above.
[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described traffic scene understanding methods.
[0010] Beneficial Effects: This application addresses the current approach, which focuses solely on understanding specific target categories and lacks a comprehensive and in-depth understanding of traffic scenarios. Therefore, after extracting image features from the current road scene image, this application sequentially obtains the detection results and trajectory information of the target object through target detection and trajectory tracking. The detection results, trajectory information, and historical scene understanding results are then input into a pre-trained scene structuring model to obtain structured data of the current road scene, thereby achieving structured prediction of the road scene and significantly improving the understanding performance of subsequent scene understanding models. Furthermore, by incorporating historical scene understanding results into the current scene structure prediction, the continuity of time can compensate for the information loss in single-frame data, significantly improving the stability of understanding in complex scenarios and solving the problem of poor adaptability of traditional techniques to sudden scenarios. Further, this application also performs target behavior recognition based on the detection results, trajectory information, and structured data to obtain the behavior recognition results of the target object. Finally, by fusing multi-dimensional information such as image features, detection results, and behavior recognition results through a scene understanding model, the current scene understanding result is output, thereby gaining a more comprehensive and in-depth understanding of the traffic scene, achieving the technical effect of improving the accuracy of traffic scene understanding and ensuring vehicle driving safety. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0012] Figure 1A flowchart of a traffic scene understanding method provided in an embodiment of this application; Figure 2 This application provides a schematic diagram of a traffic scene understanding process in accordance with an embodiment of the present application. Figure 3 A flowchart of a feature extraction method provided in an embodiment of this application; Figure 4 A schematic diagram illustrating an iterative fusion method for temporal features provided in an embodiment of this application; Figure 5 A flowchart of BEV feature fusion provided in this application embodiment; Figure 6 A schematic diagram of feature fusion based on deformable cross attention provided in an embodiment of this application; Figure 7 A diagram illustrating the process of target detection, trajectory tracking, and group target recognition provided in this application embodiment; Figure 8 A schematic diagram of a road area provided in an embodiment of this application; Figure 9 A flowchart illustrating a specific traffic scene understanding method provided in this application embodiment; Figure 10 A schematic diagram of a scene understanding model provided in an embodiment of this application; Figure 11 A schematic diagram of an information fusion module provided in an embodiment of this application; Figure 12 This is a schematic diagram of parallel processing in a processor provided in an embodiment of this application; Figure 13 This is a schematic diagram of the structure of a traffic scene understanding device provided in an embodiment of this application; Figure 14 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0014] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] The specific application environment architecture or hardware architecture upon which the traffic scene understanding method relies is described herein. Specifically, this application is particularly applicable to autonomous driving scenarios, and the hardware architecture includes sensors, a communication unit, and an edge computing unit. The sensors are specifically camera sensors, such as vehicle-mounted cameras, which may include monocular / dual-lens cameras, surround-view cameras, etc., deployed around the vehicle to collect external environmental information, such as RGB color and grayscale images of the road scene. The vehicle-mounted cameras then transmit the acquired images to the edge computing unit via the communication unit, where low-latency data transmission can be achieved through in-vehicle Ethernet. Finally, the edge computing unit processes the received images to obtain the scene understanding result of the road scene.
[0017] See Figure 1 and Figure 2 As shown, an embodiment of this application provides a traffic scene understanding method, which includes: Step S11: Extract features from the current road scene image acquired by the sensor to obtain image features, and perform target detection on the image features to obtain the detection result of the target object.
[0018] In this embodiment, an on-board edge computing unit can be configured in the intelligent driving vehicle to receive road scene images collected by the camera sensor and process the road scene images.
[0019] Specifically, after receiving the current road scene image acquired by the sensor, feature extraction is first performed on the current road scene image to obtain image features. It should be noted that the current road scene image is a sequential multi-frame image, where the specific number of frames can be set according to different application scenarios, such as 3-5 frames, 5-10 frames, or even 20 frames, etc.
[0020] In specific implementation methods, such as Figure 3As shown, feature extraction is performed on the current road scene image acquired by the sensor to obtain image features. Specifically, this includes the following steps: Step S111: Based on the preset feature extraction network, perform feature extraction on the multiple frames of time-series images acquired by the sensor to obtain the time-series features corresponding to each frame of time-series image.
[0021] In this embodiment, a preset feature extraction network is first used to extract features from multiple frames of temporal images to obtain the temporal features corresponding to each frame of temporal images. The feature extraction network can be ResNet, SwinTransformer, or other networks.
[0022] Step S112: Perform bird's-eye view transformation on each time series feature to obtain bird's-eye view features, and fuse each bird's-eye view feature to obtain fused image features.
[0023] In this embodiment, after obtaining the temporal features, the two-dimensional features of each frame in the temporal features need to be converted into bird's-eye view (BEV) features. Specifically, this process can employ a depth prediction method based on Lift-Splat-Shoot (LSS) and a query-based method for bird's-eye view conversion. Taking the LSS-based depth prediction method as an example, this method achieves the conversion from 2D features to 3D space by explicitly estimating the depth distribution of pixels in the image. First, a depth estimation network is used to predict the discrete depth distribution of pixels. Then, a 3D camera frustum feature is constructed through outer product operations. Finally, the camera frustum point cloud is assigned to the corresponding BEV grid by combining the camera's intrinsic and extrinsic parameters, and the BEV features are obtained through pooling operations.
[0024] Understandably, because camera sensors are distributed around the vehicle, they capture images from different perspectives. For example, the front-view camera sees the front of the vehicle, the side-view cameras see the sides, and the rear-view camera sees the rear. In image coordinates, it's difficult to piece together this information from different perspectives into a complete model of the surrounding environment. However, a BEV (Battery Electric Vehicle) provides a top-down view, where features from all perspectives are projected onto the same plane coordinate system directly above the vehicle. This allows information from the front, rear, left, and right cameras to be naturally aligned and stitched together, creating a 360° all-around perception of the vehicle's surroundings. Therefore, through bird's-eye view transformation, image features from different perspectives can be uniformly converted into a common, perspective-free coordinate system with physical scale, facilitating subsequent calculations.
[0025] After obtaining the bird's-eye view features of each frame of the temporal image, it is necessary to fuse these features to obtain the fused image features. In a specific implementation, fusing the bird's-eye view features to obtain the fused image features includes: determining the current target bird's-eye view feature to be fused according to the temporal sequence, and fusing the target bird's-eye view feature with historical fused image features to obtain the current fused image feature; determining the next bird's-eye view feature to be fused according to the temporal sequence, using the next bird's-eye view feature as the target bird's-eye view feature, and using the current fused image feature as the historical fused image feature, then returning to the step of fusing the target bird's-eye view feature with the historical fused image features to obtain the current fused image feature, until the fusion process of all bird's-eye view features is completed to obtain the fused image features. That is, this application discloses an iterative fusion method for temporal features, such as... Figure 4 As shown, for the BEV features of two adjacent frames, such as the BEV features at time t-1 and the BEV features at time t, feature fusion is first performed to obtain the fused image features at time t. Then, it is fused with the BEV features at time t+1 to obtain the fused image features at time t+1. This process is repeated until the temporally fused image features are finally obtained.
[0026] Specifically, when fusing the target bird's-eye view features with historical fused image features to obtain the current fused image features, the process includes: using a sparse optical flow estimation algorithm or a deep learning-based optical flow estimation algorithm to estimate the optical flow of the target bird's-eye view features and historical fused image features, thereby calculating the corresponding optical flow map; and then fusing the target bird's-eye view features, historical fused image features, and the optical flow map to obtain the current fused image features. That is, in each iteration of fusion, this application first inputs the target bird's-eye view features and historical fused image features into the optical flow estimation module to perform optical flow estimation to obtain the optical flow map, and then fuses the obtained optical flow map with the target bird's-eye view features and historical fused image features to obtain the current fused image features. The optical flow estimation module can employ a sparse optical flow estimation algorithm or a deep learning-based optical flow estimation algorithm to perform optical flow estimation, thereby obtaining the optical flow map. FlowEST (Flow Estimation) represents the optical flow estimation function, where B is the offset between two adjacent feature maps in the horizontal and vertical directions. t Let BEV be the characteristic of time t, B t-1 The BEV characteristics are at time t-1.
[0027] It should be noted that while bird's-eye view features can provide a bird's-eye perspective of traffic scenes, showing the location and shape of targets, their description of the target's motion state is relatively limited, and they are easily affected by changes in lighting and occlusion in practical applications. Optical flow maps, on the other hand, accurately capture object motion state information and can track the trajectory of fast-moving objects well. Therefore, optical flow maps are used to supplement bird's-eye view features, directly providing a reference for object motion and enhancing the understanding of dynamic changes in the scene. Furthermore, optical flow maps are two-dimensional, describing the motion vector of each pixel on the bird's-eye view features at two adjacent time points. Compared to three-dimensional representations, this method requires less data and subsequent computation, saving storage and computing resources.
[0028] Specifically, such as Figure 5 As shown, when fusing the BEV features at time t and time t-1 of adjacent time points, the BEV features at time t and time t-1 are first input into the optical flow estimation module for optical flow estimation to obtain an optical flow map. Then, the obtained optical flow map is fused with the BEV features at time t and time t-1 to obtain the fused image features at time t.
[0029] In one specific implementation, an optimization-based sparse optical flow estimation algorithm can be used for optical flow estimation. Taking the Lucas-Kanade (LK) algorithm as an example, the goal of this algorithm is to find an optimal displacement vector for the point (x, y). This minimizes the difference between the surrounding area and the two frames: ; in, This represents the optical flow vector calculated at position (x, y), where x and y are the corresponding grid coordinates on the BEV feature map. This represents the eigenvalue located at coordinate (i, j) on the BEV feature map at time t-1. This indicates that at time t, the BEV feature map is located at... Eigenvalues of coordinates For local neighborhood.
[0030] In another specific implementation, a deep learning-based optical flow estimation algorithm can be used to estimate the optical flow of the target bird's-eye view features and the historical fused image features to calculate the corresponding optical flow map. Specifically, this includes: calculating an initial optical flow based on the target bird's-eye view features and the historical fused image features, and using the initial optical flow to deform the target bird's-eye view features to obtain deformed bird's-eye view features; calculating an intermediate optical flow based on the deformed bird's-eye view features, the historical fused image features, and the initial optical flow, and using the intermediate optical flow to deform the deformed bird's-eye view features to obtain new deformed bird's-eye view features; then calculating a new intermediate optical flow based on the deformed bird's-eye view features, the historical fused image features, and the intermediate optical flow; determining whether a preset number of iterations has been reached; if not, reverting to the step of deforming the deformed bird's-eye view features using the intermediate optical flow, until the preset number of iterations is reached, and then outputting the final optical flow; and using preset weighting coefficients to perform a weighted calculation of the initial optical flow and the final optical flow to obtain the optical flow map.
[0031] When using deep learning-based optical flow estimation, the specific steps are as follows: First, the initial optical flow is calculated based on the BEV feature maps of two adjacent frames. Typically, this involves a series of convolutional and deconvolutional operations; FlowNetC is an optical flow estimation network. Then, the initial optical flow is used to estimate B. t Perform deformation to obtain the deformed bird's-eye view features, and then based on the deformed bird's-eye view features and B... t-1 The intermediate optical flow is calculated from the initial optical flow, and the intermediate optical flow is used to deform the features of the deformed bird's-eye view to obtain new deformed bird's-eye view features. Then, based on the deformed bird's-eye view features and B... t-1 Calculate the new intermediate optical flow using the intermediate optical flow.
[0032] In other words, this application employs a multi-layer optical flow network structure, which requires the use of intermediate optical flow generated in the previous iteration. BEV feature map at time t t Perform a warp operation, B t Relative deformation The warp operation is implemented using bilinear interpolation. Assuming there are k layers in total, the optical flow estimate for the k-th layer is... Thus obtaining B t Pixel positions after feature map deformation ,in, and It is optical flow Displacement in the x and y directions.
[0033] Finally, the initial optical flow and the final optical flow are weighted using preset weighting coefficients to obtain the optical flow map, i.e. The weight All are constants between [0,1], used to adjust the importance of optical flow estimation at each stage. A staged supervision method is adopted to minimize the mean square error between the optical flow estimation at each stage and the actual optical flow, thereby optimizing and obtaining better optical flow estimation performance.
[0034] Furthermore, after outputting the final optical flow, the method further includes: processing the final optical flow using a target convolutional network to obtain a processed optical flow; wherein the number of layers in the target convolutional network is greater than a preset layer threshold, and the kernel size of the target convolutional network is less than a preset size threshold; correspondingly, the initial optical flow and the final optical flow are weighted using preset weight coefficients to obtain an optical flow map, including: weighting the initial optical flow, the final optical flow, and the processed optical flow using preset weight coefficients to obtain an optical flow map.
[0035] That is, this application can also introduce a target convolutional network specifically for handling small displacement optimization, using deeper network layers and smaller convolutional kernels to process the features after multi-layer network processing. Processing is performed to obtain the processed optical flow, which is a more accurate small-displacement optical flow. Correspondingly, in the final weighting, the initial optical flow, the final optical flow obtained from the multilayer network, and the small-displacement optical flow are weighted and summed together to obtain the final optical flow map: Among them, weight It is a constant between [0,1].
[0036] In a specific implementation, the target bird's-eye view features, historical fused image features, and optical flow graph are fused to obtain the current fused image features. This includes: adding a first positional encoding corresponding to the query vector to the target bird's-eye view features to obtain initial query features; obtaining optical flow graph features obtained by linearly transforming the optical flow graph using a preset linear layer; adding optical flow graph features to the initial query features to obtain target query features; adding a second positional encoding corresponding to the value vector to the historical fused image features to obtain target value features; and fusing the target query features, target value features, and target key features based on a cross-attention mechanism to obtain the current fused image features. The target key features are either the target query features or the target value features.
[0037] It is understood that this embodiment specifically uses deformable cross-attention fusion to fuse the BEV features of two adjacent frames, with the BEV feature B at time t being... t As the query vector, the BEV feature B at time t-1 t-1 As a value vector. Specifically, as follows: Figure 6 As shown, for the BEV feature at time t, the first positional encoding corresponding to the query vector is added. To obtain the initial query features Where x and y are the grid coordinates of the BEV feature map. Further, it is necessary to obtain the optical flow map features obtained after linearly transforming the optical flow map using the optical flow map mapping unit. The optical flow map mapping unit can be implemented using a linear layer, and then optical flow map features are added to the initial query features to obtain the target query features. In this layer, the linear layer has one input channel and the number of output channels is the same as the number of BEV feature channels at time t. Furthermore, for the BEV features at time t-1, a second positional encoding corresponding to the value vector is added. To obtain the target value features The first positional code corresponding to the query vector and the second positional code corresponding to the value vector can be either fixed-rule positional codes or learnable positional codes.
[0038] It should be noted that this application improves the positional encoding of the query vector by utilizing optical flow graphs, thereby enhancing the ability of fused features to model temporal motion trends. This provides rich information input for understanding road traffic scenes and predicting the motion intentions of traffic participants. Compared to schemes that only add positional encoding to the query vector Q, the additional optical flow graph features can introduce dynamic motion priors, compensate for the shortcomings of positional encoding in modeling temporal motion trends, mitigate the impact of feature content ambiguity / occlusion, handle situations with object occlusion or sensor noise in rainy or snowy weather, improve robustness in complex dynamic scenes, and supplement subsequent tasks such as target detection, tracking, and trajectory prediction.
[0039] For key features, this embodiment provides selection methods for different scenarios. To ensure proper calculation of attention weights, the key features can be set to be the same as the value features, i.e., using the feature reuse of the previous time step t-1 (K=V), to better capture the temporal dependency between the current time and the previous time step, such as the movement trends and state changes of road traffic participants. Furthermore, if the primary purpose is to use the feature structure of the current time step to guide the sampling of features from the previous time step, the key features can be set to be the same as the query features, i.e., using the feature reuse of the current time step t (K=Q). Finally, the target query features, target value features, and target key features are input into the deformable cross-attention module for fusion to obtain the fused image features at time step t. .
[0040] After extracting the fused image features, this application further performs target detection on the image features to obtain the detection results of the target objects. It should be noted that the detection results specifically include a first detection result for a single target and a second detection result for a group of targets; a group of targets is a group composed of at least two single targets of the same type. It is understandable that, due to the large number of targets in urban traffic scenarios, such as vehicles, pedestrians, and non-motorized vehicles, the identification of individual targets and their motion intentions are easily affected by the surrounding environment, leading to information confusion. For example, at busy intersections, the trajectories of various traffic participants are mixed, highly conflicting, and severely occluded. Identifying each target individually may ignore the interaction relationships between them, and judging the motion intention based solely on the change in the motion state of a single target has a high degree of uncertainty. Therefore, by identifying and subsequently processing group targets of the same type, similar size, and similar location, the occlusion and re-identification problems of individual targets can be significantly alleviated, information confusion in interactive scenarios can be minimized, and the motion intention of targets can be better analyzed through the spatial relationships and motion trends of members within the group of targets. Therefore, this application requires the separate detection of single targets and group targets. This application completely separates the single target detection and group target detection tasks and processes them using different branches. This allows each model to focus on the identification of a specific category of targets in each branch process, which significantly reduces the training difficulty and effectively improves the recognition accuracy.
[0041] In a specific implementation, target detection is performed on image features to obtain a first detection result, including: inputting the image features to a first target detection head to obtain a first detection result output by the first target detection head, including the target bounding box, target category, and target category confidence; the target bounding box is used to determine the target identifier, target location, target size, and target orientation. That is, this application receives the fused image features through a target sub-branch k (e.g., pedestrians) and passes them through the first target detection head. Output the following content, which is the first detection result: ; Where, N k Let be the total number of targets detected by branch k; i represents the target identifier, i.e., the i-th target; The coordinates of the target bounding box can be used to determine information such as target identification, target location, target size, and target orientation.
[0042] In a specific implementation, target detection is performed on image features to obtain a second detection result, including: inputting the image features and the first detection result into a second target detection head to obtain a second detection result output by the second target detection head, which includes a group bounding box, a group member index, a group category, and a group category confidence level; wherein, the group member index is used to record the target identifier corresponding to each member in the group.
[0043] That is, this application receives the fused image features and the first target detection result through a group sub-branch g (such as a pedestrian group), and outputs the following content through the second target detection head, namely the second target detection result: ; in, This indicates that the targets included in the group must come from the corresponding category of targets in the target branch. For example, members of the "pedestrian group" must be the outputs of the "pedestrian" sub-branch in the target branch k.
[0044] Furthermore, the first loss function corresponding to the first object detection head is constructed based on the first classification loss function and the first localization loss function; the second loss function corresponding to the second object detection head is constructed based on the second classification loss function, the second localization loss function, and the member matching loss function. Specifically, the first classification loss function is constructed based on the true object category and the object category confidence; the first localization loss function is constructed based on the target bounding box output by the first object detection head and the pre-labeled true object bounding boxes; the second classification loss function is constructed based on the true group category and the group category confidence; the second localization loss function is constructed based on the group bounding boxes output by the second object detection head and the pre-labeled true group boxes; and the member matching loss function is constructed based on the target bounding box and the group bounding box.
[0045] Specifically, the first loss function includes a first classification loss function (ensuring correct category) and a first localization loss function (ensuring accurate bounding boxes). The first loss function is defined as follows: ; in, These are the weighting coefficients; The first classification loss function is the cross-entropy loss function. The first localization loss function is the GIOU (Generalized Intersection over Union) loss function.
[0046] ; In the formula, Indicates the true category of the target. Confidence level for the target category When the value is 0, it indicates the road background.
[0047] ; In the formula, For the pre-labeled real target bounding box, This is the target bounding box output by the first target detection head.
[0048] Furthermore, the second loss function includes a second classification loss function, a second localization loss function, and a member matching loss function (ensuring that the group correctly includes its members). The second loss function is defined as follows: ; in, , These are the weighting coefficients. This is the second classification loss function, which is specifically constructed based on the true group category and the group category confidence. Its definition is the same as the first classification loss function, and it is used to ensure that the group category is correct. The second localization loss function is specifically constructed based on the group bounding boxes output by the second object detection head and the pre-annotated real group boxes. Its definition is the same as the second localization loss function, and it is used to ensure the accuracy of the overall bounding boxes of the group. Match a loss function to the member to force the member target to be within the group box: ; Where, m= IOU (Intersection over Union) is used to calculate the intersection-over-union ratio.
[0049] Accordingly, the above method also includes: weighting the first loss function and the second loss function based on preset weight coefficients to construct a total loss function; and using the total loss function to jointly train the first object detection head and the second object detection head to obtain trained first and second object detection heads. That is, in this application, the total loss of the entire multi-branch model is a weighted sum of the losses of each branch, ensuring collaborative optimization among the branches. The total loss function is: ; Where w k w g Branch weights can be set according to the difficulty of the task; for example, group recognition is more difficult and can be assigned higher weights.
[0050] Furthermore, when performing target detection on image features to obtain the second detection result, the first detection result can also be input into a clustering model. The clustering model then clusters the targets based on the comparison between the distance between each target's center point and a preset distance threshold, and the second detection result is obtained based on the clustering results. That is, this embodiment can also cluster single targets based on the target's category, size, and location information in the first detection result to obtain a group of targets. For cases where the target distribution is relatively uniform, the simple and efficient K-means algorithm is sufficient. For cases where the group distribution is irregular and the degree of clustering varies, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm can discover clusters of arbitrary shapes and is more robust to data with noisy points. In addition, analysis can be performed based on different clustering levels to provide richer clustering results. Depending on the actual application requirements, multiple methods can be integrated into the software algorithm library and options can be set. During use, one or more methods can be called according to the characteristics of the traffic scenario, real-time requirements, etc., to improve the adaptability and scalability of the target group recognition method. Meanwhile, a distance threshold is used to control the density of clusters. The distance threshold is defined on the BEV plane. The density of clusters is adjusted by comparing the distance between target center points with the distance threshold.
[0051] The process of determining the preset distance threshold includes: obtaining a preset initial distance threshold and determining the traffic flow density based on the original road scene image; if the traffic flow density is greater than the preset traffic flow density, the preset initial distance threshold is reduced; if the traffic flow density is not greater than the preset traffic flow density, the preset initial distance threshold is increased. That is, the distance threshold can be adaptively adjusted. In scenarios with high traffic flow density, the distance threshold can be appropriately reduced to make the clustering more compact, so as to accurately identify closely adjacent groups of targets; conversely, when the traffic flow density is low, the distance threshold can be appropriately increased to include relatively dispersed targets in the same group, improving the inclusiveness of the clustering.
[0052] For example, for motor vehicles, traffic flow can be estimated by calculating the number of vehicles passing through per minute using Q=(60×O×S) / (L×V), where lane occupancy is O (dimensionless), average vehicle length is L (m), road segment length is S (m), and average vehicle speed is V (m / s). These values can be obtained through processing by roadside cameras, edge computing devices, etc. For pedestrians, [the calculation can be based on...]. Calculate the number of people passing through per minute to estimate pedestrian flow, including pedestrian density. (person / m) 2The value of the detection area is W (m), which is the number of pedestrians per unit area and the average walking speed of pedestrians v (m / s). The width of the detection area is W (m). All values can be obtained by processing cameras, edge computing devices, etc. installed on the road.
[0053] Furthermore, detection results that meet preset filtering conditions in the first and second detection results can be eliminated. Specifically, this includes: eliminating detection results in the first detection results where the target category confidence is less than a preset confidence threshold, and eliminating detection results in the second detection results where the group category confidence is less than a preset confidence threshold; determining whether the number of members in each group in the second detection results is at least two, and eliminating detection results where the number of members is not at least two; if any single target exists in at least two different groups, then the cross-union ratio (CUNR) between any single target and at least two different groups is calculated based on the cross-union ratio method, and the group corresponding to the largest CUNR is retained, while the remaining groups are eliminated.
[0054] It is understood that this embodiment can remove detection results in the first detection result where the confidence level of the target category is less than a preset confidence threshold, and remove detection results in the second detection result where the confidence level of the group category is less than a preset confidence threshold. For example, results with a confidence level below 0.5 are removed, and only detection results with a confidence level above 0.5 are retained. It should be noted that when the second detection result is the output of a clustering model, only detection results in the first detection result where the confidence level of the target category is less than 0.5 need to be removed. Furthermore, this embodiment must ensure the logical consistency between single targets and group targets. For example, a group must contain at least two single targets of the same category, and a single target cannot belong to multiple groups simultaneously. For example, for group g, its members are checked. If all targets originate from the corresponding category within the target branch and have at least two members, cases where a single target is mistakenly identified as a group are excluded. If a target is contained within multiple groups, the group with the highest IoU with that target is retained to eliminate conflicts, using the following formula: ; Where m represents the conflicting target. For the candidate group containing m.
[0055] Step S12: Track and identify the trajectory of the target object based on the detection results to obtain the trajectory information of the target object.
[0056] In this embodiment, the target detection results of historical frames and the target detection results of the current frame can also be used to track and identify the trajectory of the target object in order to obtain the trajectory information of the target object. The target tracking trajectory construction can employ target tracking algorithms such as DeepSORT and ByteTrack. The above processes of target detection, trajectory tracking, and group target recognition can be described as follows: Figure 7 As shown in the image.
[0057] Step S13: Input the detection results, trajectory information and historical scene understanding results into the pre-trained scene structuring model to obtain the structured data of the current road scene; the scene structuring model is a model used to perform structuring processing on the data of the road scene.
[0058] In this embodiment, the detection results, trajectory information, and historical scene understanding results are input into a pre-trained scene structuring model to obtain structured data of the current road scene, thereby achieving structured prediction of the road scene and significantly improving the understanding effect of the subsequent scene understanding model. Furthermore, by incorporating historical scene understanding results into the current scene structure prediction, the temporal continuity can be used to compensate for the information loss in single-frame data, significantly improving the stability of understanding in complex scenes and solving the problem of poor adaptability of traditional technologies to sudden scenarios.
[0059] The specific processing steps of the scene structuring model include: fusing detection results, trajectory information, and scene understanding results corresponding to historical road scene images based on a spatial attention mechanism to obtain fused features; determining the bird's-eye view corresponding to the current road scene image, and determining the road type probability distribution of each grid in the bird's-eye view based on the fused features; determining the road type of each grid based on the road type probability distribution to construct a road type map, and extracting the boundaries of different road types from the road type map to obtain the road regions of the corresponding road types; and obtaining the structured data of the current road scene based on the positional relationship between each road region and the vehicle. In other words, this application fuses detection results, trajectory information, and historical scene understanding results based on a spatial attention mechanism, enabling the fused features to simultaneously capture static regions and dynamic targets, overcoming the shortcomings of fragmented modeling in traditional methods. Furthermore, by introducing historical scene understanding results into the fusion, the temporal stability of the road scene is used to compensate for the information loss in single-frame data, significantly improving the understanding stability in complex scenes. Further, this application calculates the road type probability distribution of each grid in the bird's-eye view based on the obtained fused features, and determines the road type of the grid through the probability distribution. Compared with the hard classification of traditional semantic segmentation, this can handle the problem of blurred region boundaries more delicately, improving the accuracy of the road type map. Ultimately, by constructing structured data of the current road scenario based on the positional relationship between the road area and the vehicle, it can directly adapt to the decision-making needs of autonomous driving, reduce the conversion costs of downstream systems, improve decision-making efficiency, and ensure the safety and accuracy of decision-making in complex traffic scenarios.
[0060] The method involves fusing detection results, trajectory information, and scene understanding results corresponding to historical road scene images based on a spatial attention mechanism to obtain fused features. This includes: encoding the detection results, trajectory information, and scene understanding results corresponding to historical road scene images separately to obtain corresponding feature encoding results; concatenating the feature encoding results to obtain a feature concatenation matrix; calculating an attention weight matrix based on the spatial attention mechanism and the feature concatenation matrix; and weighting the feature concatenation matrix using the attention weight matrix to obtain the final fused features. Specifically, this application encodes and concatenates cross-modal information such as detection results, trajectory information, and scene understanding results corresponding to historical road scene images separately to obtain a feature concatenation matrix F. Then, it calculates an attention weight matrix Attn based on the spatial attention mechanism and the feature concatenation matrix, and finally weights the feature concatenation matrix using the attention weight matrix to obtain the final fused features, thereby enhancing road structure-related information. ; in, For element-wise multiplication, This is the final fusion feature.
[0061] In a specific implementation, the attention weight matrix is calculated based on the spatial attention mechanism and the feature concatenation matrix, including: performing global average pooling on the feature concatenation matrix to obtain a global feature vector; performing a two-dimensional convolution operation on the global feature vector to obtain the convolution result; and processing the convolution result using a first preset activation function to obtain the attention weight matrix. The attention weight matrix is calculated as follows: ; Where GlobalAvgPool represents the global average pooling operation, Conv2D represents the two-dimensional convolution operation, and Sigmoid is the activation function.
[0062] Furthermore, the probability distribution of road types for each grid in the bird's-eye view is determined based on the fused features, including: processing the fused features using a preset convolutional classifier to obtain the feature representation of each grid in the bird's-eye view; and processing the feature representation of each grid separately using a second preset activation function to obtain the probability distribution of road types for each grid. That is, from a bird's-eye view perspective, road information within the visible range uses the BEV map grid resolution, and a preset convolutional classifier and the preset activation function Softmax are used to predict the probability distribution of road types for each grid. ; in, These correspond to the road background, driveway, sidewalk, and pedestrian waiting area, respectively. H and W represent the number of BEV grids in the width and height directions, respectively, and i and j represent the grid indices. It can be understood that the probability distribution reflects the model's likelihood of identifying the road type of each grid. Each grid has a set of probability values corresponding to different road types; the higher the probability value, the higher the probability that the grid belongs to the corresponding road type. For example, P(10,20,1)=0.9 means that the probability of grid (10,20) belonging to a lane is 90%.
[0063] When determining the road type of each grid based on the road type probability distribution to construct a road type map, the specific steps include: for any grid, taking the road type corresponding to the maximum probability in the road type probability distribution as the road type of that grid; road types include lanes, sidewalks, and pedestrian waiting areas; assigning values to each grid using numerical labels corresponding to each road type to construct a road type map; and accordingly, extracting the boundaries of different road types from the road type map to obtain the road area of the corresponding road type, including: obtaining a set of target grids with the same numerical label from the road type map, and processing the target grid set using a preset boundary extraction algorithm to obtain the road area of the corresponding road type.
[0064] That is, after obtaining the road type probability distribution for each grid, the road type corresponding to the maximum probability is selected as the predicted road type for that grid. Road types can include road background, lanes, sidewalks, and pedestrian waiting areas. Then, each grid cell is assigned a value using a numerical label (0, 1, 2, 3) corresponding to each road type to construct a road type map. In this way, probability distributions are transformed into specific type labels, thereby achieving a structured classification of road scenes. In other words, this application uses a road type map to structurally represent road scenes at the grid level, intuitively presenting the potential structure of different areas within the road scene through different probability value distributions.
[0065] After obtaining the predicted road type map, based on the judgment results for different types, the continuous boundaries of each type of region are extracted. Specifically, this application obtains a set of target meshes with the same numerical label from the road type map, and then processes the target mesh set using a preset boundary extraction algorithm to obtain the road region corresponding to the road type. For example, for the lane region, by judging the meshes with a value of 1 in the type map, a specialized boundary extraction algorithm is used to determine the continuous boundaries of the lane region, thereby obtaining the lane region. For example: ; ; ; The above are the boundary sets of the lane, sidewalk, and pedestrian waiting area, respectively.
[0066] It should also be noted that, to ensure the accuracy and structural consistency of predictions, the loss function of the scene structuring model consists of three parts: 1. Classification Loss (Focal Loss, which addresses class imbalance): ; in For true type tags (one-hot encoded), For category weights, For focus parameters (such as) =2); 2. Boundary Consistency Loss (ensuring IoU between the predicted boundary and the true boundary): ; in To predict the boundary, For the true boundary; 3. Temporal stability loss (constraining the coherence of predictions for adjacent frames): ; in This is the road type map for frame t. The road type map is shown in frame t+1. MSE (Mean Squared Error) is the mean squared error loss function to avoid abrupt changes in road structure, such as lanes suddenly disappearing.
[0067] The total loss function is: ; in, For example, weighting coefficients .
[0068] Furthermore, structured data of the current road scene is obtained based on the positional relationship between each road area and the vehicle. This includes: obtaining a pre-defined starting position for each road area corresponding to any road type; and numbering each road area clockwise from the starting position based on the positional relationship between each road area and the vehicle to obtain structured data of the current road scene. Assuming the final obtained road areas are as follows... Figure 8 As shown, after obtaining each road area, the vehicle is used as the reference. Based on the positional relationship between each road area and the vehicle, the lane areas, pedestrian areas, and pedestrian waiting areas in different directions are numbered in a specific order to ensure consistency with the scene description.
[0069] The numbering process is as follows: Coordinate system definition: The origin (0,0) is the vehicle's position, the vehicle's orientation is forward (positive x-axis), and the right side is the positive y-axis. Pedestrian waiting area numbering: Based on the direction of the vehicle, the four pedestrian waiting areas in the diagram are numbered clockwise, starting from the left front of the vehicle, and are designated as pedestrian waiting area-1, pedestrian waiting area-2, pedestrian waiting area-3, and pedestrian waiting area-4. Lane numbering: based on intersection ( Figure 8 Using the rectangular area at the center as the boundary, the four lane areas are numbered clockwise, starting from the lane where the vehicle is located, and are designated as Lane-1, Lane-2, Lane-3, and Lane-4.
[0070] Pedestrian crossing numbering: Similar to lane numbering, the four pedestrian crossings are numbered clockwise, starting from the pedestrian crossing of the lane where the vehicle is located, and are recorded as pedestrian crossing-1, pedestrian crossing-1, pedestrian crossing-3, and pedestrian crossing-4.
[0071] Step S14: Based on the detection results, trajectory information and structured data, identify the behavior of the target object to obtain the behavior identification result of the target object.
[0072] In this embodiment, target behavior recognition is also required based on detection results, trajectory information, and structured data to obtain the behavior recognition result of the target object.
[0073] Step S15: Input the image features, detection results and behavior recognition results into the pre-trained scene understanding model to obtain the scene understanding results of the current road scene.
[0074] In this embodiment, a scene understanding model is used to fuse multi-dimensional information such as image features, detection results, and behavior recognition results to output the current scene understanding result. This allows for a more comprehensive and in-depth understanding of the traffic scene, thereby improving the accuracy of traffic scene understanding and ensuring vehicle driving safety.
[0075] The scene understanding model is built upon multiple Transformer blocks. The scene understanding result for the current road scene includes a textual description of the scene understanding and meta-features obtained after processing by multiple Transformer blocks. That is, in addition to utilizing existing language models, the scene understanding model can be built using multiple Transformer blocks. In the output layer of the scene understanding model, in addition to the textual description of the scene understanding, it also includes meta-features processed by multiple Transformer blocks for use by downstream tasks. Meta-feature vectors include [0.8 (left turn intention), 0.6 (acceleration trend), 0.3 (surrounding vehicle density)], etc.
[0076] As can be seen, this application addresses the current approach by focusing only on understanding specific target categories and lacking a comprehensive and in-depth understanding of traffic scenarios. Therefore, after extracting image features from the current road scene image, this application sequentially obtains the detection results and trajectory information of the target object through target detection and trajectory tracking. The detection results, trajectory information, and historical scene understanding results are then input into a pre-trained scene structuring model to obtain structured data of the current road scene, thereby achieving structured prediction of the road scene and significantly improving the understanding performance of subsequent scene understanding models. Furthermore, by incorporating historical scene understanding results into the current scene structure prediction, the continuity of time can compensate for the information loss in single-frame data, significantly improving the stability of understanding in complex scenarios and solving the problem of poor adaptability of traditional techniques to sudden scenarios. Further, this application also performs target behavior recognition based on the detection results, trajectory information, and structured data to obtain the behavior recognition results of the target object. Finally, by fusing multi-dimensional information such as image features, detection results, and behavior recognition results through a scene understanding model, the current scene understanding result is output, thereby gaining a more comprehensive and in-depth understanding of the traffic scene, achieving the technical effect of improving the accuracy of traffic scene understanding and ensuring vehicle driving safety.
[0077] See Figure 9 As shown, this application discloses a specific method for understanding traffic scenarios. Compared to the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically, it includes: Step S21: Extract features from the current road scene image acquired by the sensor to obtain image features, and perform target detection on the image features to obtain the detection result of the target object.
[0078] Step S22: Track and identify the trajectory of the target object based on the detection results to obtain the trajectory information of the target object.
[0079] Step S23: Input the detection results, trajectory information and historical scene understanding results into the pre-trained scene structuring model to obtain the structured data of the current road scene; the scene structuring model is a model used to perform structuring processing on the data of the road scene.
[0080] Step S24: Input the detection results and trajectory information into the temporal convolutional network to extract the motion trend features of the target object, and use the embedding layer and graph convolutional network to extract features from the structured data to obtain scene features.
[0081] This embodiment also discloses a behavior recognition model. After obtaining the detection results, trajectory information, and structured data, these data are used as input to the behavior recognition model to output the behavior recognition result of the target object. It should be noted that the behavior recognition result is a set of labels based on the road scene structure and the target trajectory definition, namely (target type: starting area - ending area). By fusing the moving target category and its movement trend, the behavior of the target object is characterized in a fine-grained manner, thereby providing high-level semantic information for scene understanding.
[0082] For example, for vehicle targets or vehicle group targets, the behavioral labels include: going straight in the same direction, going straight in opposite directions, lane 1-lane 2, lane 1-lane 3, lane 1-lane 4, lane 2-lane 1, lane 2-lane 3, lane 2-lane 4, lane 3-lane 1, lane 3-lane 2, lane 3-lane 4, lane 4-lane 1, lane 4-lane 2, lane 4-lane 3, and stopping to wait. For pedestrian targets or pedestrian groups, the behavioral tags include: pedestrian waiting area 1-pedestrian waiting area 2, pedestrian waiting area 1-pedestrian waiting area 3, pedestrian waiting area 1-pedestrian waiting area 4, pedestrian waiting area 2-pedestrian waiting area 1, pedestrian waiting area 2-pedestrian waiting area 3, pedestrian waiting area 2-pedestrian waiting area 4, pedestrian waiting area 3-pedestrian waiting area 1, pedestrian waiting-3-pedestrian waiting area 2, pedestrian waiting area 3-pedestrian waiting area 4, pedestrian waiting area 4-pedestrian waiting area 1, pedestrian waiting area 4-pedestrian waiting area 2, pedestrian waiting area 4-pedestrian waiting area 3, waiting.
[0083] Specifically, in the behavior recognition model, this embodiment employs different feature extraction networks to extract the motion trend features and scene features of the target object. First, the detection results and trajectory information are input into a temporal convolutional network (TCN) to extract the motion trend features of the target object. This is used to reflect the trajectory dynamics of the target object; simultaneously, embedding layers and graph convolutional networks (GCNs) are used to extract features from structured data to obtain scene features. This is used to reflect the relationship between the target object and the road area.
[0084] Step S25: Use motion trend features as query vectors and scene features as key vectors to calculate attention scores. Use the attention scores to perform weighted fusion of motion trend features and scene features to obtain the first fused feature. Then, use a preset prediction head to process the first fused feature to obtain the behavior recognition result of the target object.
[0085] In this embodiment, after obtaining the motion trend features and scene features, the motion trend features are further processed... Used as query vectors, and scene features Used as a key vector to calculate attention scores Then, attention scores are used to weight and fuse motion trend features and scene features to obtain the first fused feature. To highlight key information related to behavioral tags: ; in, For element-wise multiplication, concat represents the concatenation function.
[0086] Specifically, the obtained behavior recognition result is a triplet including target type, starting region, and ending region. The preset prediction head includes a target type prediction head, a starting region prediction head, and an ending region prediction head. Correspondingly, the first fusion feature is processed using the preset prediction head to obtain the behavior recognition result of the target object, including: inputting the first fusion feature into the target type prediction head to determine the target type; inputting the first fusion feature and the embedding feature corresponding to the target type into the starting region prediction head to determine the starting region; and inputting the first fusion feature, the embedding feature corresponding to the target type, and the embedding feature corresponding to the starting region into the ending region prediction head to determine the ending region.
[0087] That is, the triplet behavior label corresponding to the behavior recognition result is Where c is the target type, such as r s As the starting region, r e This is the endpoint region. Correspondingly, there are three prediction heads: a target type prediction head, a starting region prediction head, and an ending region prediction head. Therefore, when processing the first fused feature using the preset prediction heads, it is necessary to model the joint probability of these three factors. Dependencies are modeled using conditional probability chains: ; Specifically, the first fused feature is first input into the target type prediction head, and the target type prediction head is used to output the target type: ; where N c Number of types; The first fused feature and the embedded features corresponding to the target type are then input into the starting region prediction head to determine the starting region: N r Given the number of regions, the input contains embedding features of type c; Finally, the first fusion feature, the embedding feature corresponding to the target type, and the embedding feature corresponding to the starting region are input into the ending region prediction head to determine the ending region: The input contains type c and starting region r. s Embedded features.
[0088] For example, the output is: Given that the probability of a vehicle is 0.95 and the probability of a pedestrian is 0.05, we can conclude that c represents a vehicle. If the probability of lane 1 is 0.8, then r s Lane 1; If the probability of lane 2 is 0.7, then r e Lane 2; Final label: (vehicle, lane 1, lane 2), indicating that the vehicle changes lanes from lane 1 to lane 2.
[0089] Step S26: Perform feature mapping processing on the image features to obtain mapped image features; extract features from each target object in the detection results to obtain detection result features, and aggregate the detection result features to obtain aggregated detection result features; extract features from each target object in the behavior recognition results to obtain target behavior features, and aggregate the target behavior features to obtain aggregated target behavior features.
[0090] In this embodiment, as Figure 10 As shown, image features, detection results, and behavior recognition results are first processed by the information fusion module to obtain fused features, and then processed by the scene understanding model to predict the scene understanding result of the current road scene.
[0091] Furthermore, such as Figure 11 As shown, the information fusion module includes three mapping units: an image feature mapping unit, a detection result mapping unit, and a behavior result mapping unit, which are used to process image features, detection results, and behavior recognition results, respectively. Specifically, image features are input to the image feature mapping unit for processing to obtain mapped image features; detection results are the length, width, height, three-dimensional spatial coordinates, target category probability, and target orientation angle of each target object, which are processed into a vector and then processed by the detection result mapping unit to obtain detection result features. These detection result features are then aggregated to obtain aggregated detection result features; behavior recognition results are the probability distributions of each behavior category, which are processed into a feature vector and then processed by the behavior result mapping unit to obtain target behavior features with the same number of image feature channels. These target behavior features are then aggregated to obtain aggregated target behavior features.
[0092] It should also be noted that, in order to improve the model processing speed, this embodiment accelerates the processing by using parallel processing for the branches corresponding to the three mapping units in the information fusion module. Specifically, the above method further includes: determining a pre-allocated first thread pool, a second thread pool, and a third thread pool; wherein the first thread pool, the second thread pool, and the third thread pool operate in parallel mode; and allocating image features, detection results, and behavior recognition results to the first thread pool, the second thread pool, and the third thread pool respectively for parallel processing to obtain the corresponding mapped image features, aggregated detection result features, and aggregated target behavior features. That is, as shown... Figure 12 As shown, the AI processor pre-allocates three thread pools, namely thread pool 1, thread pool 2 and thread pool 3. By distributing the three branches of image features, detection results and behavior recognition results to different thread pools for parallel processing, the corresponding mapped image features, aggregated detection result features and aggregated target behavior features are obtained, thereby improving computational efficiency.
[0093] Step S27: Perform feature fusion on the mapped image features, aggregated detection result features, and aggregated target behavior features to obtain the second fused features, and input the second fused features into the pre-trained scene understanding model to obtain the scene understanding results of the current road scene.
[0094] In this embodiment, as Figure 11 As shown, the mapped image features, aggregated detection result features, and aggregated target behavior features are then fused through a feature fusion module to obtain the second fused feature. Optional feature fusion methods include concatenating features along the channel dimension, directly adding feature values element-wise, or using other feature fusion methods. Subsequently, several neural network units, such as convolutional layers or Transformer blocks, are used to process the fused features to obtain the final second fused feature. Finally, the second fused feature is input into a pre-trained scene understanding model to obtain the scene understanding result of the current road scene. The scene understanding model can be constructed using multiple Transformer blocks or directly using an existing language model. In the output layer of the scene understanding model, in addition to the textual description of the scene understanding, it also contains features processed by multiple Transformer blocks as meta-features for scene understanding and intent prediction.
[0095] The following are examples of text descriptions for scene understanding: 1. Two cars: A blue car is traveling straight from east to west, possibly heading to its next destination indicated in the straight lane; a red car is waiting for a left-turn signal on a north-to-south road, possibly because its destination requires it to turn onto the left side of the street. 2. A motorcycle: Located in the right lane of the east-to-west road, preparing to go straight when the green light turns on, possibly because going straight is the most direct route to its destination. 3. Three pedestrians: Two pedestrians are crossing the road from south to north on the crosswalk, possibly because they are going to the commercial area or bus stop on the other side; a third pedestrian is waiting on the crosswalk at the northeast corner of the intersection, possibly because he is waiting for a safe time to join the flow of people crossing the road. 4. A bicycle: Traveling from west to east on the bicycle lane on the south side of the road, because the bicycle lane provides a dedicated path for safe riding.
[0096] As can be seen, this application also discloses a behavior recognition model that comprehensively utilizes detection results, trajectory information, and structured data to identify the behavior of target objects, effectively improving the understanding effect of subsequent scene understanding models. The scene understanding model is then used to process image features, detection results, and behavior recognition results to achieve a deeper understanding of road scenes and prediction of target intentions. This results in a more comprehensive and deeper understanding of traffic scenes, improving the accuracy of traffic scene understanding and ensuring vehicle driving safety.
[0097] Furthermore, it should be noted that this application performs model training in a cloud data center, training the perception model on a dataset of collected and labeled road scene images. The model is optimized using a backpropagation algorithm, where the optimizer can employ Gradient Descent (GD), Adaptive Moment Estimation (Adam), or Adam Weight Decay Regularization (AdamW, an Adam optimization algorithm), among others. The proposed model acceleration method improves training efficiency. After model training is complete, the trained model is deployed to an edge computing platform.
[0098] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0099] See Figure 13 As shown, embodiments of this application also provide a traffic scene understanding device, which includes: The image processing module 11 is used to extract features from the current road scene image acquired by the sensor to obtain image features, and to perform target detection on the image features to obtain the detection result of the target object; The trajectory recognition module 12 is used to track and recognize the trajectory of the target object based on the detection results, so as to obtain the trajectory information of the target object; The scene structuring module 13 is used to input the detection results, trajectory information and historical scene understanding results into the pre-trained scene structuring model to obtain the structured data of the current road scene; the scene structuring model is a model used to perform structured processing on the data of the road scene. The behavior recognition module 14 is used to recognize the behavior of the target object based on the detection results, trajectory information and structured data, so as to obtain the behavior recognition result of the target object; The scene understanding module 15 is used to input image features, detection results and behavior recognition results into a pre-trained scene understanding model to obtain the scene understanding results of the current road scene.
[0100] As can be seen, this application addresses the current approach by focusing only on understanding specific target categories and lacking a comprehensive and in-depth understanding of traffic scenarios. Therefore, the image processing and trajectory recognition modules of this application extract image features from the current road scene image and then sequentially perform target detection and trajectory tracking to obtain the detection results and trajectory information of the target object. In the scene structuring module, the detection results, trajectory information, and historical scene understanding results are input into a pre-trained scene structuring model to obtain structured data of the current road scene, thereby achieving structured prediction of the road scene and significantly improving the understanding effect of subsequent scene understanding models. Furthermore, by incorporating historical scene understanding results into the current scene structure prediction, the continuity of time can compensate for the information loss in single-frame data, significantly improving the stability of understanding in complex scenarios and solving the problem of poor adaptability of traditional technologies to sudden scenarios. Further, this application also uses a behavior recognition module to perform behavior recognition of the target object based on the detection results, trajectory information, and structured data to obtain behavior recognition results. Finally, in the scene understanding module, the scene understanding model integrates multi-dimensional information such as image features, detection results, and behavior recognition results to output the current scene understanding result, thereby gaining a more comprehensive and in-depth understanding of the traffic scene, achieving the technical effect of improving the accuracy of traffic scene understanding and ensuring vehicle driving safety.
[0101] For a description of the features in the embodiment corresponding to the traffic scene understanding device, please refer to the relevant description of the embodiment corresponding to the traffic scene understanding method, which will not be repeated here.
[0102] Embodiments of this application also provide an electronic device, such as... Figure 14 As shown, it can specifically be an edge computing unit device, including a memory, a processor, a communication interface, an input / output interface, and a communication bus. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described traffic scene understanding method embodiments.
[0103] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the traffic scene understanding method when it is run.
[0104] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0105] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the traffic scene understanding method.
[0106] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the traffic scene understanding method.
[0107] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0108] The foregoing has provided a detailed description of a traffic scene understanding method, device, medium, and product provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for understanding traffic scenarios, characterized in that, include: Feature extraction is performed on the current road scene image acquired by the sensor to obtain image features, and target detection is performed on the image features to obtain the detection result of the target object; Based on the detection results, the trajectory of the target object is tracked and identified to obtain the trajectory information of the target object; The detection results, trajectory information, and historical scene understanding results are input into a pre-trained scene structuring model to obtain structured data of the current road scene; The scene structured model is a model used to structure road scene data. The behavior of the target object is identified based on the detection results, the trajectory information, and the structured data to obtain the behavior identification result of the target object; The image features, the detection results, and the behavior recognition results are input into a pre-trained scene understanding model to obtain the scene understanding results of the current road scene; The historical scene understanding result is the scene understanding result output by the scene understanding model after processing the historical road scene images before the current road scene image; the structured data is the data obtained by extracting the road regions of each road type from the road type map corresponding to the current road scene, and numbering each road region based on the positional relationship between each road region and the vehicle; the road type map is constructed based on the road type probability distribution of each grid in the bird's-eye view.
2. The traffic scene understanding method according to claim 1, characterized in that, The current road scene image is a multi-frame temporally consecutive image. Accordingly, the step of extracting features from the current road scene image acquired by the sensor to obtain image features includes: Based on a preset feature extraction network, features are extracted from multiple frames of time-series images acquired by the sensor to obtain the time-series features corresponding to each frame of time-series image; A bird's-eye view transformation is performed on each of the time-series features to obtain bird's-eye view features, and the bird's-eye view features are fused to obtain fused image features.
3. The traffic scene understanding method according to claim 2, characterized in that, The process of fusing the bird's-eye view features to obtain fused image features includes: The target bird's-eye view features to be fused are determined according to the chronological order, and the target bird's-eye view features are fused with the historical fused image features to obtain the current fused image features; The next bird's-eye view feature to be fused is determined according to the chronological order, and the next bird's-eye view feature to be fused is used as the target bird's-eye view feature. The current fused image feature is used as the historical fused image feature. Then, the process jumps back to the step of fusing the target bird's-eye view feature with the historical fused image feature to obtain the current fused image feature. This process continues until the fusion process of all bird's-eye view features is completed to obtain the fused image feature.
4. The traffic scene understanding method according to claim 3, characterized in that, The step of fusing the target bird's-eye view features with historical fused image features to obtain the current fused image features includes: Optical flow estimation is performed on the target bird's-eye view features and historical fused image features using a sparse optical flow estimation algorithm or a deep learning-based optical flow estimation algorithm to calculate the corresponding optical flow map; The target bird's-eye view features, the historical fused image features, and the optical flow map are fused to obtain the current fused image features.
5. The traffic scene understanding method according to claim 4, characterized in that, Optical flow estimation is performed on the target bird's-eye view features and historical fused image features using a deep learning-based optical flow estimation algorithm to calculate the corresponding optical flow map, including: The initial optical flow is calculated based on the target bird's-eye view features and the historical fused image features, and the target bird's-eye view features are deformed using the initial optical flow to obtain deformed bird's-eye view features. The intermediate optical flow is calculated based on the deformed bird's-eye view features, the historical fused image features, and the initial optical flow. The deformed bird's-eye view features are then deformed using the intermediate optical flow to obtain new deformed bird's-eye view features. Finally, the new intermediate optical flow is calculated based on the deformed bird's-eye view features, the historical fused image features, and the intermediate optical flow. Determine whether the preset number of iterations has been reached. If not, jump back to the step of deforming the features of the deformed bird's-eye view using intermediate optical flow until the preset number of iterations is reached, and then output the final optical flow. The initial optical flow and the final optical flow are weighted using preset weighting coefficients to obtain an optical flow map.
6. The traffic scene understanding method according to claim 5, characterized in that, Following the final output optical flow, the following is also included: The final optical flow is processed using a target convolutional network to obtain the processed optical flow; wherein the number of layers in the target convolutional network is greater than a preset layer threshold, and the kernel size of the target convolutional network is smaller than a preset size threshold. Accordingly, the step of weighting the initial optical flow and the final optical flow using preset weighting coefficients to obtain an optical flow map includes: The initial optical flow, the final optical flow, and the processed optical flow are weighted using preset weighting coefficients to obtain an optical flow map.
7. The traffic scene understanding method according to claim 4, characterized in that, The process of fusing the target bird's-eye view features, the historical fused image features, and the optical flow map to obtain the current fused image features includes: Add a first position code corresponding to the query vector to the target bird's-eye view features to obtain initial query features, and obtain the optical flow map features obtained by linearly transforming the optical flow map using a preset linear layer, so as to add the optical flow map features to the initial query features to obtain target query features; Add a second positional encoding corresponding to the value vector to the historical fused image features to obtain the target value features; The target query feature, the target value feature, and the target key feature are fused based on the cross-attention mechanism to obtain the current fused image feature; wherein, the target key feature is either the target query feature or the target value feature.
8. The traffic scene understanding method according to claim 1, characterized in that, The step of identifying the behavior of the target object based on the detection results, the trajectory information, and the structured data to obtain the behavior identification result of the target object includes: The detection results and trajectory information are input into a temporal convolutional network to extract the motion trend features of the target object, and the embedded layer and graph convolutional network are used to extract features from the structured data to obtain scene features. The motion trend feature is used as the query vector, and the scene feature is used as the key vector to calculate the attention score. The attention score is then used to perform a weighted fusion of the motion trend feature and the scene feature to obtain the first fused feature. The first fused feature is processed using a preset prediction head to obtain the behavior recognition result of the target object.
9. The traffic scene understanding method according to claim 8, characterized in that, The behavior recognition result is a triplet including target type, start region and end region, and the preset prediction head includes target type prediction head, start region prediction head and end region prediction head; Accordingly, the step of processing the first fused features using a preset prediction head to obtain the behavior recognition result of the target object includes: The first fused feature is input into the target type prediction head to determine the target type; The first fusion feature and the embedding feature corresponding to the target type are input into the starting region prediction head to determine the starting region; The first fusion feature, the embedding feature corresponding to the target type, and the embedding feature corresponding to the starting region are input into the ending region prediction head to determine the ending region.
10. The traffic scene understanding method according to claim 1, characterized in that, The step of inputting the image features, the detection results, and the behavior recognition results into a pre-trained scene understanding model to obtain the scene understanding results of the current road scene includes: The image features are subjected to feature mapping processing to obtain mapped image features; For each target object in the detection results, feature extraction is performed to obtain detection result features, and the detection result features are aggregated to obtain aggregated detection result features; For each target object in the behavior recognition result, feature extraction is performed to obtain target behavior features, and the target behavior features are aggregated to obtain aggregated target behavior features; The mapped image features, the aggregated detection result features, and the aggregated target behavior features are fused to obtain a second fused feature; The second fused feature is input into a pre-trained scene understanding model to obtain the scene understanding result of the current road scene.
11. The traffic scene understanding method according to claim 10, characterized in that, Also includes: A first thread pool, a second thread pool, and a third thread pool are pre-allocated; wherein the first thread pool, the second thread pool, and the third thread pool operate in a parallel mode. The image features, the detection results, and the behavior recognition results are respectively assigned to the first thread pool, the second thread pool, and the third thread pool for parallel processing to obtain the corresponding mapped image features, the aggregated detection result features, and the aggregated target behavior features.
12. The traffic scene understanding method according to any one of claims 1 to 11, characterized in that, The scene understanding model is built based on multiple Transformer blocks, and the scene understanding result of the current road scene includes a scene understanding text description and meta-features obtained after processing by the multiple Transformer blocks.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the traffic scene understanding method as described in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the traffic scene understanding method as described in any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the traffic scene understanding method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Traffic scene understanding method and device based on video streaming
CN112347933A
Method and vehicle with an advanced driver assistance system for risk-based traffic scene analysis
US20150344030A1