Reighting all-in-one machine traffic target real-time identification system based on edge calculation

By using an edge computing-based radar-visual integrated system, combined with millimeter-wave radar and traffic monitoring cameras, multi-source data fusion and intelligent scheduling are achieved at roadside terminals. This solves the problems of recognition delay and false detection in complex environments of existing systems, and improves the real-time performance and robustness of traffic target recognition.

CN121789154AInactive Publication Date: 2026-04-03HEFEI VISION WAVE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing intelligent traffic monitoring systems struggle to achieve real-time and accurate identification of multiple types of traffic targets in complex environments. They also have a high dependence on network bandwidth and central computing power, resulting in high identification delays and false positive/false negative rates, failing to meet the real-time and robustness requirements of complex traffic scenarios.

Method used

The system adopts a real-time traffic target recognition system based on edge computing and radar-visual integration. Through data fusion between millimeter-wave radar and traffic monitoring cameras, it completes multi-source perception, area division, intelligent scheduling, target recognition and structured reporting on the roadside embedded terminal. It introduces radar-guided spatial attention, multi-event target decoupling branches and radar-visual dynamic fusion, and dynamically adjusts the task inference frame rate and computing resource allocation.

Benefits of technology

Balancing recognition accuracy, end-to-end latency, and bandwidth usage in complex traffic environments, and reducing reliance on central servers, this system achieves real-time and reliable traffic target recognition and event determination, making it suitable for large-scale deployment in complex road environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789154A_ABST
    Figure CN121789154A_ABST
Patent Text Reader

Abstract

The invention discloses a thunder-vision all-in-one machine traffic target real-time identification system based on edge calculation, and the system comprises a perception preprocessing module which is used for generating a radar guide map and a preprocessing video frame; the region dividing and buffering module is used for dividing and buffering regions of interest to form a multi-source sensing data sequence; the edge scheduling control module is used for setting a task set and generating a scheduling strategy for each task; the improved BiSeNet recognition module is used for constructing an improved BiSeNet network and outputting a traffic target recognition result and a traffic event feature map; the result generation module is used for forming structured traffic target data and event data; and the storage and report module is used for storing the data in a local medium and uploading the data to the central platform through the communication module. According to the invention, high-precision and low-bandwidth real-time identification of multiple types of traffic targets is realized through Leiyu fusion and improved BiSeNet identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to a real-time traffic target recognition system based on edge computing and radar-visual integrated machine. Background Technology

[0002] Most existing intelligent traffic monitoring systems employ video surveillance, loop detectors, single millimeter-wave radar, or microwave detection equipment to perform simple data collection and triggering at the roadside, while target detection, behavior analysis, and event recognition are completed at a central platform. Front-end devices typically only perform encoding compression and basic filtering; large amounts of raw video streams or semi-structured data are transmitted back to the cloud or data center via wired or wireless networks, where a centralized server performs deep learning inference and multi-algorithm combinations. In urban arterial roads, expressways, and highways with numerous lanes, high traffic volume, and frequently changing scenarios, the current model is highly dependent on network bandwidth and central computing power. Once the link becomes congested or the server load increases, end-to-end processing latency increases significantly, making it difficult to guarantee the real-time performance of illegal behavior identification and traffic anomaly warnings. In complex environments such as rain, snow, fog, haze, low light at night, backlight, and strong glare, the perception capability of a single camera or radar is limited, making it difficult to reliably identify multiple types of traffic targets, including pedestrians, motor vehicles, non-motorized vehicles, and debris, easily leading to missed detections and false detections, thus affecting traffic safety management effectiveness.

[0003] Some existing systems are beginning to deploy embedded intelligent all-in-one machines on the roadside, introducing object detection networks and simple rule engines. However, these systems typically lack fine-grained task scheduling mechanisms for end-to-end latency, and cannot dynamically adjust inference frame rates, computing power allocation, and task priorities based on the real-time complexity of traffic scenarios. This can easily lead to algorithm congestion or delayed recognition results during peak hours. Traditional semantic segmentation or object detection networks are mostly designed for single image inputs, failing to fully utilize the complementary characteristics of radar point clouds and video images. They lack radar-guided spatial attention, multi-event target decoupling output, and a dynamic radar-visual fusion structure, making it difficult to achieve integrated recognition of traffic target categories, behavioral states, and event risks at the network edge. Existing roadside equipment outputs mostly scattered detection results or simple counting data, lacking structured traffic target data and structured event data in a unified format. This makes it difficult to balance recognition accuracy, processing latency, and bandwidth usage under limited computing power and unstable network conditions, limiting the comprehensive analysis and coordinated control of the backend platform.

[0004] Therefore, how to provide a real-time traffic target recognition system based on edge computing and radar vision is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a real-time traffic target recognition system based on edge computing and integrated radar-visual technology. This invention achieves full-process processing—from multi-source perception, region segmentation, intelligent scheduling, target recognition, to structured reporting—on a roadside embedded terminal through data fusion between millimeter-wave radar and traffic monitoring cameras. The invention constructs a perception preprocessing module, a region segmentation and buffering module, an edge scheduling control module, an improved BiSeNet recognition module, a result generation module, and a storage and reporting module. It utilizes a multi-level inference scheduling unit under latency constraints and an adaptive computing power allocation mechanism based on the dynamic complexity of traffic scenarios to dynamically adjust the inference frame rate and computing resources of each task at the edge based on the number of traffic targets, lane occupancy, and environmental conditions. At the recognition network level, it introduces radar-guided spatial attention, multi-event target decoupling branches, and a radar-visual dynamic fusion head to achieve integrated output of traffic target categories, behavioral states, and event risks. Compared to existing single-sensor plus cloud-based centralized processing solutions, this invention can balance recognition accuracy, end-to-end latency, and bandwidth usage in complex traffic environments, making it suitable for large-scale deployment at road intersections, urban arterial roads, and highways.

[0006] A real-time traffic target recognition system based on edge computing and radar-visual integrated machine according to an embodiment of the present invention includes the following modules: The perception preprocessing module is used to collect radar echo data and traffic scene video data and preprocess them to generate radar guidance maps and preprocessed video frames. The region segmentation and buffering module is used to segment and buffer regions of interest based on radar guidance maps and pre-processed video frames to form multi-source sensing data sequences. The edge scheduling control module is used to set up task sets and generate scheduling strategies for each task. An improved BiSeNet recognition module is used to construct an improved BiSeNet network and output traffic target recognition results and traffic event feature maps; The results generation module is used to generate structured traffic target data and event data based on the scheduling strategy, traffic target identification results, and traffic event feature maps. The storage and reporting module is used to store structured traffic target data and event data, as well as corresponding key video frames, on local media and upload them to the central platform through the communication module.

[0007] Optionally, modules can be integrated using the following methods: The device is initialized in the edge computing terminal, and radar echo data and on-site traffic scene video data from the radar vision integrated machine are collected, preprocessed respectively, and radar guidance map and preprocessed video frames are generated. Based on the road area and lane model, the radar guidance map and pre-processed video frames are divided into regions of interest and buffered to form a multi-source sensing data sequence arranged in chronological order. A real-time intelligent processing module is constructed at the edge, and a task set is set for radar data analysis, traffic event identification and data uploading. An end-to-end latency threshold is set for each task using a multi-level inference scheduling unit, and a traffic complexity index is calculated through an adaptive computing power allocation mechanism to obtain the scheduling strategy. An improved BiSeNet network was constructed, which spatially weighted the multi-source perception data sequence by radar-guided spatial attention, introduced a multi-event target decoupling branch for traffic feature extraction, and performed cross-modal fusion in the radar-visual dynamic fusion head to obtain traffic target recognition results and traffic event feature maps. Based on the scheduling strategy, traffic target trajectories, lane occupancy information, and traffic event judgment results are generated according to the traffic target identification results and traffic event feature maps, forming structured traffic target data and event data; Structured traffic target data and event data, along with corresponding key video frames, are stored on local media and uploaded to the central platform via a communication module.

[0008] Optionally, the initialization equipment includes a millimeter-wave radar of the radar vision integrated machine, a traffic monitoring camera, and a communication module.

[0009] Optionally, generating the radar guidance map and preprocessed video frames includes: In the edge computing terminal, the millimeter-wave radar, traffic monitoring camera and communication module are powered on and started. The working mode, transmission power, sampling frequency and clock synchronization parameters of the millimeter-wave radar and the resolution, exposure time, frame rate and clock synchronization parameters of the traffic monitoring camera are initialized and configured so that the millimeter-wave radar and traffic monitoring camera work based on a unified time reference. The millimeter-wave radar is controlled to collect raw radar echo data covering the road monitoring area according to the scanning cycle, and encapsulated into continuous radar frame data according to a unified timestamp. The traffic monitoring camera is controlled to collect traffic scene video images according to the frame rate, and the collected image frames are marked according to the timestamp corresponding to the radar frame data to form a video frame sequence that is aligned with the radar frame data in time. The radar frame data is filtered, corrected, and target point cloud is extracted. Based on the calibration parameters stored in the edge computing terminal, the three-dimensional coordinates of the target points are mapped to image plane coordinates to generate a radar guidance map that is spatially aligned with the video frame sequence. The video frame sequence is then subjected to size scaling, color space conversion, and pixel value normalization to generate preprocessed video frames corresponding to the radar guidance map.

[0010] Optionally, forming a multi-source sensing data sequence arranged in chronological order includes: The configured road area contour parameters and lane model parameters are read in the edge computing terminal. The lane model parameters include the lane center line, lane boundary line and lane number information. The road area contour parameters and lane model parameters are loaded into the scene of the region of interest division. The radar guidance map and pre-processed video frames are mapped to the coordinate system corresponding to the scene. The road monitoring area is extracted based on the road area contour parameters. Based on the lane model parameters, the corresponding lane interest area and roadside interest area are divided on the radar guidance map and pre-processed video frames, and area identifiers are assigned to each interest area. According to the acquisition time sequence, the radar guidance map after region of interest division and the pre-processed video frame are paired. Data with the same timestamp and region identifier are combined into a frame of multi-source sensing data. Multiple consecutive frames of multi-source sensing data are stored in the buffer in time order to form a multi-source sensing data sequence arranged in time order.

[0011] Optionally, obtaining the scheduling strategy includes: In the edge-side real-time intelligent processing module, a task set is created, which includes radar data analysis tasks, traffic event recognition tasks, and data upload tasks. Input data type and output data type are defined for each task. Each task is assigned an end-to-end latency threshold by a multi-level inference scheduling unit. The end-to-end latency threshold includes the maximum allowable processing time from the entry of multi-source sensing data into the edge computing terminal to the generation of the task output result. The task is divided into a multi-level priority queue according to the real-time requirements of different tasks, and a multi-level inference scheduling table is generated. Through an adaptive computing power allocation mechanism, the number of traffic targets, lane occupancy ratio, and environmental state parameters within the current time window are statistically analyzed from the multi-source perception data sequence. The traffic complexity index is calculated by combining the number of traffic targets, lane occupancy ratio, and environmental state parameters according to preset weight coefficients. The traffic complexity index is equal to the number of traffic targets multiplied by the first weight coefficient, plus the lane occupancy ratio multiplied by the second weight coefficient, plus the environmental state parameters multiplied by the third weight coefficient. Based on the traffic complexity index and the multi-level inference scheduling table, the inference frame rate, processing time slice length, and allocated computing power of radar data analysis tasks, traffic event identification tasks, and data upload tasks are dynamically adjusted to generate a scheduling strategy that includes task execution order, task execution cycle, and resource allocation ratio.

[0012] Optionally, obtaining the traffic target recognition result and the traffic event feature map includes: Preprocessed video frames are read from multi-source sensing data sequences and input into the spatial and contextual paths of the improved BiSeNet network to obtain video feature maps at different scales. At the same time, the radar guidance map at the corresponding time is aligned to the size of the video feature map in terms of spatial scale. Radar guidance spatial attention calculation is performed on the radar guidance map. The attention weight coefficient of each spatial location is calculated, and the attention weight coefficient is multiplied with the video feature map at the corresponding spatial location to generate a video feature map that has undergone radar guidance spatial weighting. The video feature map, which has undergone radar-guided spatial weighting, is fused and compressed to generate a visual semantic feature map. The radar-guided map is then subjected to convolutional transformation and scale alignment to generate a radar semantic feature map. The visual semantic feature map and the radar semantic feature map are input into the radar-visual dynamic fusion head. The fusion weight is calculated based on the radar semantic feature map, and the visual semantic feature map is weighted to obtain a cross-modal fusion feature map. The cross-modal fusion feature map is input into the multi-event target decoupling branch, and the target basic feature map is output through convolutional layers and classification layers. The first decoupling sub-branch outputs the traffic target category feature map based on the target basic feature map. The second decoupling sub-branch uses the traffic target category feature map as a query to perform attention filtering on the target basic feature map and outputs the traffic target behavior state feature map. The third decoupling branch uses the traffic target category feature map and the traffic target behavior state feature map as context to perform risk attention modeling on the target basic feature map and outputs the traffic event risk feature map. By combining traffic target category feature maps, traffic target behavior state feature maps, and traffic event risk feature maps with cross-modal fusion feature maps, pixel-level or region-level decisions are made to generate traffic target recognition results and traffic event feature maps.

[0013] Optionally, the formation of structured traffic target data and event data includes: Pixel-level target masks are extracted from traffic target recognition results and traffic event feature maps. Pixel sets with the same target category label and spatial connectivity are clustered into traffic target instances. A frame-level target list containing target number, target category, pixel region and timestamp is generated. Range, speed and azimuth information are extracted from radar data with the same timestamp. The frame-level target list is matched with radar range, velocity, and azimuth according to timestamp and spatial location. Range, velocity, and azimuth values ​​are added to each traffic target instance. Target instances with similar spatial locations and the same target category in adjacent frames are associated in chronological order to generate continuous target trajectories. The continuous target trajectory includes target number, trajectory coordinate sequence, velocity sequence, and time sequence. Based on the lane model parameters, the trajectory coordinates are mapped to the lane coordinate system. The occupancy time of each lane is accumulated within the statistical time, and the lane occupancy ratio is calculated. The lane occupancy ratio is equal to the cumulative lane occupancy time divided by the statistical time length. Based on the comparison of speed sequence with speed limit threshold, comparison of travel direction with lane direction, and trajectory dwell time, speeding, wrong-way driving, long-term lane occupation, and littering events are determined. The target trajectory, lane occupancy ratio, and event records are encapsulated into structured traffic target data and structured traffic event data.

[0014] Optionally, storing structured traffic target data and event data, along with corresponding key video frames, on a local medium and uploading them to the central platform via a communication module includes: In the edge computing terminal, a unique index identifier is assigned to the structured traffic target data, structured traffic event data and key video frames. The index identifier is combined with the timestamp, device number and lane number to generate a storage file name or record number, and the data is written to the cache area in the local storage medium. The data in the buffer is stored and managed in a cyclical manner according to time order. When the space occupied by the buffer exceeds the limit, the earliest data record is deleted, and only the structured traffic target data, structured traffic event data and key video frames in the most recent time window are retained. Structured traffic target data, structured traffic event data, and key video frames are read from the cache according to the upload cycle and priority order, packaged into upload data frames, and sent to the central platform via wireless link.

[0015] The beneficial effects of this invention are: This invention achieves local acquisition, preprocessing, and spatiotemporal alignment of traffic target information at the edge terminal by constructing functional modules on the roadside for perception preprocessing, area segmentation and buffering, edge scheduling control, result generation, and storage and reporting. This is achieved in conjunction with multi-source fusion of millimeter-wave radar and traffic monitoring cameras. The edge scheduling control module introduces a multi-level inference scheduling unit under latency constraints and an adaptive computing power allocation mechanism based on the dynamic complexity of traffic scenarios. Under different traffic flow, lane occupancy, and environmental conditions, it dynamically adjusts the priority, inference frame rate, and resource allocation ratio of tasks such as radar data analysis, traffic event recognition, and data uploading, ensuring stable operation of key recognition links within end-to-end latency thresholds. Compared with existing solutions that rely on a single sensor and centralized cloud processing, this invention significantly reduces the dependence on central server computing power and backhaul bandwidth. It can maintain real-time perception and continuous operation capabilities even in scenarios with network fluctuations or bandwidth limitations, improving the system's practicality and reliability in complex road environments.

[0016] This invention improves upon BiSeNet's network structure by introducing radar-guided spatial attention, multi-event target decoupling branches, and a radar-visual dynamic fusion head. It deeply integrates radar-guided maps with video features, providing unified output based on traffic target categories, behavioral states, and event risks. Continuous target trajectories are generated through pixel-level clustering and multi-frame correlation. Combined with lane models, it calculates lane occupancy ratios and identifies various traffic events such as speeding, wrong-way driving, prolonged lane occupancy, and littering, forming structured traffic target and event data in a unified format. Furthermore, a local caching and hierarchical uploading mechanism for key video frames enables efficient linkage with the central platform and intuitive backtracking of abnormal scenarios. Compared to existing technologies that only output partial detection results or simple statistics, this invention exhibits higher recognition accuracy and robustness under complex weather conditions, varying lighting, and occlusion scenarios, facilitating intelligent control and refined operation management of road intersections, urban arterial roads, and highways. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0018] Figure 1 This is a flowchart of a real-time traffic target recognition system based on edge computing and integrated radar-visual system proposed in this invention; Figure 2 This is a structural block diagram of a real-time traffic target recognition method based on edge computing and radar-visual integrated machine proposed in this invention; Figure 3 This is a functional diagram of the improved BiSeNet network for a real-time traffic target recognition method based on edge computing in a radar-visual integrated machine proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1 A real-time traffic target recognition system based on edge computing and radar-visual integrated machine includes the following modules: The perception preprocessing module is used to collect radar echo data and traffic scene video data and preprocess them to generate radar guidance maps and preprocessed video frames. The region segmentation and buffering module is used to segment and buffer regions of interest based on radar guidance maps and pre-processed video frames to form multi-source sensing data sequences. The edge scheduling control module is used to set up task sets and generate scheduling strategies for each task. An improved BiSeNet recognition module is used to construct an improved BiSeNet network and output traffic target recognition results and traffic event feature maps; The results generation module is used to generate structured traffic target data and event data based on the scheduling strategy, traffic target identification results, and traffic event feature maps. The storage and reporting module is used to store structured traffic target data and event data, as well as corresponding key video frames, on local media and upload them to the central platform through the communication module.

[0021] refer to Figure 2 and Figure 3 A real-time traffic target recognition method based on edge computing and radar-visual integrated machine includes: The device is initialized in the edge computing terminal, and radar echo data and on-site traffic scene video data from the radar vision integrated machine are collected, preprocessed respectively, and radar guidance map and preprocessed video frames are generated. Based on the road area and lane model, the radar guidance map and pre-processed video frames are divided into regions of interest and buffered to form a multi-source sensing data sequence arranged in chronological order. A real-time intelligent processing module is constructed at the edge, and a task set is set for radar data analysis, traffic event identification and data uploading. An end-to-end latency threshold is set for each task using a multi-level inference scheduling unit, and a traffic complexity index is calculated through an adaptive computing power allocation mechanism to obtain the scheduling strategy. An improved BiSeNet network was constructed, which spatially weighted the multi-source perception data sequence by radar-guided spatial attention, introduced a multi-event target decoupling branch for traffic feature extraction, and performed cross-modal fusion in the radar-visual dynamic fusion head to obtain traffic target recognition results and traffic event feature maps. Based on the scheduling strategy, traffic target trajectories, lane occupancy information, and traffic event judgment results are generated according to the traffic target identification results and traffic event feature maps, forming structured traffic target data and event data; Structured traffic target data and event data, along with corresponding key video frames, are stored on local media and uploaded to the central platform via a communication module.

[0022] In this embodiment, the initialization device includes a millimeter-wave radar of the radar vision integrated machine, a traffic monitoring camera, and a communication module.

[0023] In this embodiment, generating the radar guidance map and preprocessed video frames includes: In the edge computing terminal, the millimeter-wave radar, traffic monitoring camera and communication module are powered on and started. The working mode, transmission power, sampling frequency and clock synchronization parameters of the millimeter-wave radar and the resolution, exposure time, frame rate and clock synchronization parameters of the traffic monitoring camera are initialized and configured so that the millimeter-wave radar and traffic monitoring camera work based on a unified time reference. The millimeter-wave radar is controlled to collect raw radar echo data covering the road monitoring area according to the scanning cycle, and encapsulated into continuous radar frame data according to a unified timestamp. The traffic monitoring camera is controlled to collect traffic scene video images according to the frame rate, and the collected image frames are marked according to the timestamp corresponding to the radar frame data to form a video frame sequence that is aligned with the radar frame data in time. The radar frame data undergoes filtering, correction, and target point cloud extraction. Based on calibration parameters stored in the edge computing terminal, the three-dimensional coordinates of the target points are mapped to image plane coordinates to generate a radar guidance map spatially aligned with the video frame sequence. The video frame sequence is then subjected to scaling, color space conversion, and pixel value normalization to generate preprocessed video frames corresponding to the radar guidance map. Specifically, the process of mapping the three-dimensional coordinates of the target points to image plane coordinates based on calibration parameters stored in the edge computing terminal to generate a radar guidance map spatially aligned with the video frame sequence involves: The edge computing terminal stores the calibration parameters of the radar and the camera. The calibration parameters include two parts: one part is the external calibration parameters, which are used to describe the spatial attitude relationship between the radar coordinate system and the camera coordinate system, including the three-dimensional rotation relationship and the three-dimensional translation relationship; the other part is the internal calibration parameters of the camera, which are used to describe the geometric characteristics of the camera imaging, including the horizontal focal length, the vertical focal length, the principal point position of the imaging plane, and the lens distortion parameters. When processing radar frame data, a single target point is extracted from the target point cloud. For each target point, the measurement results in polar coordinate form are converted into three-dimensional spatial coordinates with the radar as the origin based on the distance, horizontal angle and elevation angle measured by the radar. This gives the spatial position of the target point in the radar coordinate system. Using external calibration parameters, this three-dimensional spatial position is transformed from the radar coordinate system to the camera coordinate system. This completes one spatial coordinate transformation from the radar coordinate system to the camera coordinate system for each target point. Projection operations are performed using the camera's internal calibration parameters to project the three-dimensional spatial position onto the camera's imaging plane. The row and column coordinates of the target point in the image are calculated, and distortion correction is performed on the row and column coordinates using distortion parameters to obtain the pixel position consistent with the actual video frame. For target points with positive distance and pixel positions falling within the effective range of the image, encoded values ​​are written to the corresponding positions in a two-dimensional array with the same resolution as the video frame. The encoded values ​​are either distance information or reflection intensity information. After completing the projection and writing operations for all target points in the same frame, a radar guidance map spatially aligned with the current video frame is obtained.

[0024] In this embodiment, forming a multi-source sensing data sequence arranged in chronological order includes: The configured road area contour parameters and lane model parameters are read in the edge computing terminal. The lane model parameters include the lane center line, lane boundary line and lane number information. The road area contour parameters and lane model parameters are loaded into the scene of the region of interest division. The radar guidance map and preprocessed video frames are mapped to the coordinate system corresponding to the scene. The road monitoring area is extracted based on the road area contour parameters. Based on the lane model parameters, the corresponding lane interest regions and roadside interest regions are divided on the radar guidance map and preprocessed video frames, and area identifiers are assigned to each interest region. Specifically, the division of the corresponding lane interest regions and roadside interest regions based on the lane model parameters on the radar guidance map and preprocessed video frames involves: By using scene calibration relationships, the lane centerline point series and boundary line point series are transformed from the scene coordinate system to the image coordinate system of the radar guidance map and pre-processed video frames, and the centerline and left and right boundary contours of each lane are drawn in the image. The pixel area between the left and right boundary lines of the same lane is defined as the lane interest region, and a unique number is assigned to each lane interest region as the region identifier. The remaining valid area within the road monitoring area that does not belong to any lane interest region is divided into roadside interest regions. According to the acquisition time sequence, the radar guidance map after region of interest division and the pre-processed video frame are paired. Data with the same timestamp and region identifier are combined into a frame of multi-source sensing data. Multiple consecutive frames of multi-source sensing data are stored in the buffer in time order to form a multi-source sensing data sequence arranged in time order.

[0025] In this embodiment, obtaining the scheduling strategy includes: In the edge-side real-time intelligent processing module, a task set is created, which includes radar data analysis tasks, traffic event recognition tasks, and data upload tasks. Input data type and output data type are defined for each task. Each task is assigned an end-to-end latency threshold using a multi-level inference scheduling unit. This threshold includes the maximum allowable processing time from the entry of multi-source sensing data into the edge computing terminal to the generation of the task's output result. Furthermore, tasks are divided into multi-level priority queues based on their real-time requirements, generating a multi-level inference scheduling table. Specifically, the division of tasks into multi-level priority queues based on their real-time requirements is as follows: The multi-level inference scheduling unit configures real-time parameters for each type of task, including the maximum allowable processing latency, the expected update cycle, and the level of impact of the task on traffic safety, classifying tasks into three categories: strong real-time tasks, general real-time tasks, and non-real-time tasks. Multiple priority queues are established in the order of strong real-time, general real-time and non-real-time. Strong real-time tasks are added to the highest priority queue, general real-time tasks are added to the middle priority queue, and non-real-time tasks are added to the lowest priority queue. The execution cycle, time slice length and preemption rules of each type of task are recorded in the queue. The runtime scheduling unit selects executable tasks in order of priority from high to low according to the priority queue. Through an adaptive computing power allocation mechanism, the number of traffic targets, lane occupancy ratio, and environmental state parameters within the current time window are statistically analyzed from multi-source sensing data sequences. The traffic complexity index is then calculated by combining these three parameters according to preset weighting coefficients. The traffic complexity index equals the number of traffic targets multiplied by a first weighting coefficient, plus the lane occupancy ratio multiplied by a second weighting coefficient, plus the environmental state parameters multiplied by a third weighting coefficient. Specifically, the preset weights are: The traffic target quantity, lane occupancy ratio, and environmental status parameters obtained within the current time window are normalized. Each of the three indicators is divided by its maximum reference value in the historical sample so that all three indicators fall within the range of zero to one. Based on the traffic management department's emphasis on road operation safety, the influence weight of the traffic target quantity is set to be greater than that of the lane occupancy ratio, and the influence weight of the lane occupancy ratio is set to be greater than that of the environmental status parameters. Under the specified constraints, the first weight coefficient is set to 0.5, the second weight coefficient is set to 0.3, and the third weight coefficient is set to 0.2, so that the sum of the three weights equals one. Based on the traffic complexity index and the multi-level inference scheduling table, the inference frame rate, processing time slice length, and allocated computing power for radar data analysis, traffic incident identification, and data upload tasks are dynamically adjusted to generate a scheduling strategy that includes task execution order, task execution cycle, and resource allocation ratio. Specifically, the scheduling strategy is as follows: The scheduling strategy is a set of task scheduling parameters generated by the multi-level inference scheduling unit to constrain the execution mode of radar data analysis tasks, traffic event identification tasks and data upload tasks. The scheduling strategy gives the priority level, inference frame rate, single processing time slice length and the proportion of processor computing power that can be used for each type of task, and determines the number of executions and execution order within a scheduling cycle. When the traffic complexity index increases, the scheduling strategy sets the traffic event identification task and radar data analysis task as high priority, increases the inference frame rate and computing power ratio of the two types of tasks, extends the processing time slice, and at the same time reduces the execution frequency and computing power ratio of the data upload task. When the traffic complexity index decreases, the scheduling strategy correspondingly reduces the inference frame rate of the high priority task, freeing up computing power for data upload and non-real-time tasks.

[0026] In this embodiment, obtaining the traffic target recognition result and the traffic event feature map includes: Preprocessed video frames are read from a multi-source sensing data sequence and input into the spatial and contextual paths of an improved BiSeNet network to obtain video feature maps at different scales. Simultaneously, the radar guidance map at the corresponding time is spatially aligned to the size of the video feature maps. Specifically, obtaining video feature maps at different scales involves: In the spatial path, multiple convolutional layers and downsampling operations are passed sequentially to reduce the size of the feature map proportionally, while edge information and local detail information are extracted layer by layer to obtain a set of spatial feature maps with progressively decreasing resolution. A structure with multiple downsampling is adopted to perform multi-stage feature extraction on preprocessed video frames, and output multi-level semantic feature maps at different downsampling depths. Each level of feature map corresponds to a different receptive field and semantic abstraction level. Radar guidance spatial attention calculation is performed on the radar guidance map, calculating the attention weight coefficient for each spatial location. This attention weight coefficient is then multiplied by the video feature map at the corresponding spatial location to generate a video feature map that has undergone radar guidance spatial weighting. Specifically, the calculation of the attention weight coefficient for each spatial location involves: Based on the position of each pixel in the radar guidance map, the radar intensity value and distance information of the current position are used as input. The radar intensity value and distance information of the current position are smoothed and normalized within a local window centered on the current position to obtain the response value that reflects the salience of the local target. Within the entire radar guidance map, the response value of all positions is scaled to the range between zero and one to represent the spatial attention of the current position. The degree of attention at each pixel location is used as the attention weight coefficient of the location. When the target echo at a certain location in the radar guidance map is strong or the probability of the target being present is high, the corresponding attention weight coefficient is larger. When a certain location in the radar guidance map is close to the background or has no effective target, the corresponding attention weight coefficient is smaller. The video feature map, after radar-guided spatial weighting, undergoes feature fusion and channel compression to generate a visual semantic feature map. The radar-guided map is then convolved and scale-aligned to generate a radar semantic feature map. Both the visual and radar semantic feature maps are input into the radar-visual dynamic fusion head. Fusion weights are calculated based on the radar semantic feature map, and the visual semantic feature map is then weighted to obtain a cross-modal fusion feature map. Where: The feature fusion and channel compression are performed as follows: Feature maps at different scales are upsampled to a uniform spatial size and then concatenated in the channel direction to overlay detailed and semantic information at different scales to form a fused video feature map. The number of channels is compressed by using convolution and non-linear activation operations. Specifically, multiple convolution kernels are used to perform linear transformations in the channel direction, mapping the original large number of feature channels to fewer feature channels. At the same time, edge information, texture information and semantic information that are useful for traffic target recognition are preserved, resulting in a visual semantic feature map with fewer channels and more concentrated information. The calculation of the fusion weight is specifically as follows: the radar dynamic fusion head performs statistical analysis on the radar semantic feature map in the spatial dimension to obtain response values ​​that reflect the importance of each spatial location or channel. Then, these response values ​​are normalized, and the normalized response values ​​are used as the fusion weights. The cross-modal fused feature map is input into the multi-event target decoupling branch. Through convolutional and classification layers, it outputs a target base feature map. The first decoupling sub-branch outputs a traffic target category feature map based on the target base feature map. The second decoupling sub-branch uses the traffic target category feature map as a query to perform attention filtering on the target base feature map and outputs a traffic target behavior state feature map. The third decoupling sub-branch uses the traffic target category feature map and the traffic target behavior state feature map as context to perform risk attention modeling on the target base feature map and outputs a traffic event risk feature map. Where: The output of the traffic target category feature map based on the target basic feature map is specifically as follows: several convolutional layers and non-linear activation layers are connected in series on the target basic feature map to extract high-level semantic features that are more sensitive to category discrimination. A classification head is connected in the channel dimension to map the channel response of each spatial location to the response value of each traffic target category. The response of each category is normalized to obtain a traffic target category feature map that represents the category confidence on a pixel-by-pixel basis in the image space. The process of using the traffic target category feature map as a query to perform attention filtering on the target basic feature map and then outputting the traffic target behavior state feature map is as follows: the category feature map is compressed into a single-channel attention map in the channel direction through a one-to-one convolution layer. The values ​​of the entire attention map are normalized by summation so that the sum of the attention values ​​at all positions is one, resulting in an attention weight map. All channels at each pixel position in the target basic feature map are multiplied by the attention weight of the corresponding position to achieve category-guided feature enhancement. The weighted features are then fed into the convolutional layer and the classification layer to output a traffic target behavior state feature map that represents the probability of different behavior states pixel by pixel. The method of using traffic target category feature map and traffic target behavior state feature map as context to perform risk attention modeling on the target basic feature map is as follows: concatenating them into a context feature map in the channel dimension, then calculating the risk response value of each pixel position through multiple convolutions, and normalizing it by summation in the entire response map to obtain a risk attention weight map. The risk attention weight map is multiplied onto the target basic feature map position by position, and the features of areas with high risk probability are amplified. Finally, through convolutional layers and classification layers, the weighted features are mapped to a pixel-by-pixel traffic event risk probability distribution, i.e., a traffic event risk feature map. By combining traffic target category feature maps, traffic target behavior state feature maps, and traffic event risk feature maps with cross-modal fusion feature maps, pixel-level or region-level decisions are made to generate traffic target recognition results and traffic event feature maps.

[0027] In this embodiment, the formation of structured traffic target data and event data includes: Pixel-level target masks are extracted from traffic target recognition results and traffic event feature maps. Pixel sets with the same target category label and spatial connectivity are clustered into traffic target instances, generating a frame-level target list containing target number, target category, pixel region, and timestamp. Range, velocity, and azimuth information are extracted from radar data with the same timestamp. Specifically, the extraction of pixel-level target masks involves: The category prediction map of each frame is read from the traffic target recognition results. The category prediction map assigns a target category label to each pixel, including pedestrian category, motor vehicle category, non-motor vehicle category, litter category and background category. The category prediction map is traversed pixel by pixel. Pixels whose category label does not belong to the background are marked as foreground pixels, and pixels whose category label belongs to the background are marked as background pixels, thus obtaining the initial distribution of foreground pixels and background pixels. Foreground pixels are grouped according to target category. Connected component labeling is used for foreground pixels in the same category. Spatially connected pixel sets are found, and each spatially connected pixel set is defined as a pixel-level target mask. The frame-level target list is matched with radar range, velocity, and azimuth according to timestamp and spatial location. Range, velocity, and azimuth values ​​are added to each traffic target instance. Target instances with similar spatial locations and the same target category in adjacent frames are associated in chronological order to generate continuous target trajectories. The continuous target trajectory includes target number, trajectory coordinate sequence, velocity sequence, and time sequence. Based on the lane model parameters, the trajectory coordinates are mapped to the lane coordinate system. The occupancy time of each lane is accumulated within a statistical period, and the lane occupancy ratio is calculated. The lane occupancy ratio equals the accumulated lane occupancy time divided by the statistical time length. Speeding, driving against traffic, prolonged lane occupation, and littering events are determined based on a comparison of the speed sequence with the speed limit threshold, a comparison of the travel direction with the lane direction, and the trajectory dwell time. The target trajectory, lane occupancy ratio, and event records are encapsulated into structured traffic target data and structured traffic event data. The speed limit threshold is specifically: The speed limit for each lane is read from the road design documents or speed limit information issued by the traffic management department, and the value is written into the speed limit configuration table of the edge computing terminal along with the corresponding lane number. For each target trajectory, the speed limit of the lane is read from the speed limit configuration table according to the lane to which the trajectory belongs. The speed limit is multiplied by one and a fixed overspeed tolerance coefficient of 0.1 is added. The result is used as the speed limit threshold of the lane.

[0028] In this embodiment, storing structured traffic target data and event data, along with corresponding key video frames, on a local medium and uploading them to the central platform via a communication module includes: In the edge computing terminal, a unique index identifier is assigned to the structured traffic target data, structured traffic event data and key video frames. The index identifier is combined with the timestamp, device number and lane number to generate a storage file name or record number, and the data is written to the cache area in the local storage medium. The data in the buffer is stored and managed in a cyclical manner according to time order. When the space occupied by the buffer exceeds the limit, the earliest data record is deleted, and only the structured traffic target data, structured traffic event data and key video frames in the most recent time window are retained. Structured traffic target data, structured traffic event data, and key video frames are read from the cache according to the upload cycle and priority order, packaged into upload data frames, and sent to the central platform via wireless link.

[0029] Example 1: To verify the feasibility of this invention in practice, it was applied to a complex interchange at the intersection of a third ring road and an airport expressway. The interchange spans approximately one kilometer, encompassing six main lanes and four ramps, with an average daily traffic volume of about 120,000 vehicles. During morning and evening rush hours, the average speed ranges from 30 to 40 kilometers per hour, with heavy vehicles accounting for nearly 20%. The existing system used a single camera with centralized cloud-based recognition; the front end only uploaded the video stream, while the central computer room ran the algorithm. After long-term operation, traffic management departments found that nighttime accidents were frequent, warnings were lacking in rainy and foggy weather, and during peak hours, there were issues with missed detection of traffic targets, delayed detection of wrong-way driving, and incidents involving littering. Many anomalies could only be confirmed manually afterward, failing to provide continuous and reliable data support for rapid on-site response.

[0030] In the current interchange scenario, the system of this invention is deployed at three key diversion points. Each point is equipped with a millimeter-wave radar from a 77 GHz radar-vision integrated machine and a 2-megapixel traffic monitoring camera. Edge computing terminals are installed in roadside cabinets and connected to the traffic management network via fiber optic cables. During system operation, the millimeter-wave radar and camera in the radar-vision integrated machine synchronously collect data. The perception preprocessing module generates a radar guidance map aligned with the video and preprocessed video frames. The region segmentation and buffering module divides the regions of interest by lane and ramp and forms a multi-source perception data sequence. The edge scheduling and control module dynamically adjusts the inference frame rate and computing power allocation for traffic event recognition and data upload based on traffic complexity. The improved BiSeNet network completes the integrated recognition of traffic target categories, behavioral states, and event risks. The result generation module outputs target trajectories, lane occupancy information, and event records of speeding, wrong-way driving, prolonged lane occupation, and littering. The storage and reporting module uploads structured data in real time when the network is normal, caches it locally when the network is interrupted, and re-uploads it after network recovery.

[0031] During a period of continuous field operation and verification, the system stably outputs a large amount of structured traffic target data and traffic event data, continuously covering various typical operating conditions such as morning and evening rush hours, ordinary weekdays, weekends, and inclement weather. Comparison with manual sampling results shows that the traffic event identification results maintain high consistency under complex weather and nighttime conditions, and the alarm response for key lanes is significantly faster than the original solution. Even with limited network bandwidth or short-term interruptions, the front end can maintain uninterrupted identification and local recording, and can automatically retransmit key event records after communication is restored. Overall operational results indicate that this invention significantly improves upon the problems of insufficient identification accuracy, excessive response latency, and high bandwidth consumption in complex interchange scenarios, providing a more reliable data foundation for traffic management departments to conduct refined control and rapid response.

[0032] Table 1 Performance Comparison of Different Solutions in Complex Interchange Scenarios

[0033] As shown in Table 1, the recognition accuracy of the system of the present invention is 96%, and the event false negative rate is 3%, which is the highest and lowest among all solutions. The accuracy of single-camera cloud is only 86% with a false negative rate of 11%, the accuracy of single radar cloud is 81% with a false negative rate of 15%, the accuracy of front-end lightweight vision is 88% with a false negative rate of 9%, the accuracy of radar-vision cloud inference is 91% with a false negative rate of 7%, and the accuracy of general edge box is 92% with a false negative rate of 6%. It can be seen that the system of the present invention minimizes the false negative rate while ensuring high accuracy, making it more suitable for practical deployment on important road sections.

[0034] In terms of latency and bandwidth usage, the average latency of the system in this invention is 95 milliseconds, significantly lower than the 320 milliseconds for a single-camera cloud, 290 milliseconds for a single-radar cloud, 260 milliseconds for radar cloud inference, and 210 milliseconds for lightweight front-end vision, only slightly higher than the 150 milliseconds for a general edge box. Regarding peak bandwidth, the system in this invention achieves only 3 megabits per second, compared to 14 megabits per second for a single-camera cloud, 12 megabits per second for radar cloud inference, 10 megabits per second for a single-radar cloud, 8 megabits per second for lightweight front-end vision, and 6 megabits per second for a general edge box. This invention achieves both latency control within the hundreds of milliseconds range and the lowest bandwidth usage among all solutions.

[0035] In terms of adaptability to complex environments and system stability, the system of this invention achieves an accuracy rate of 93% in rainy and foggy weather and 94% at night, both higher than the 72% and 74% of single-camera cloud, 78% and 70% of single-radar cloud, 75% and 80% of front-end lightweight vision, 84% and 85% of radar-cloud inference, and 86% and 87% of general edge boxes. Regarding the duration of interruptions, the system of this invention can work continuously for 65 minutes, higher than the 5 minutes of single-camera cloud, 8 minutes of single-radar cloud, 20 minutes of front-end lightweight vision, 15 minutes of radar-cloud inference, and 40 minutes of general edge boxes. Its comprehensive score reaches 9.3 points, significantly better than the 6.2, 5.8, 7.1, 7.8, and 8.1 points of other solutions, demonstrating its advantages in robustness in complex environments and overall performance.

[0036] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A real-time traffic target recognition system based on edge computing and integrated radar-visualization, characterized in that, Includes the following modules: The perception preprocessing module is used to collect radar echo data and traffic scene video data and preprocess them to generate radar guidance maps and preprocessed video frames. The region segmentation and buffering module is used to segment and buffer regions of interest based on radar guidance maps and pre-processed video frames to form multi-source sensing data sequences. The edge scheduling control module is used to set up task sets and generate scheduling strategies for each task. An improved BiSeNet recognition module is used to construct an improved BiSeNet network and output traffic target recognition results and traffic event feature maps; The results generation module is used to generate structured traffic target data and event data based on the scheduling strategy, traffic target identification results, and traffic event feature maps. The storage and reporting module is used to store structured traffic target data and event data, as well as corresponding key video frames, on local media and upload them to the central platform through the communication module.

2. The method for real-time traffic target recognition based on edge computing in a radar-view integrated machine according to claim 1, applied to the real-time traffic target recognition system based on edge computing in a radar-view integrated machine according to claim 1, characterized in that, include: The device is initialized in the edge computing terminal, and radar echo data and on-site traffic scene video data from the radar vision integrated machine are collected, preprocessed respectively, and radar guidance map and preprocessed video frames are generated. Based on the road area and lane model, the radar guidance map and pre-processed video frames are divided into regions of interest and buffered to form a multi-source sensing data sequence arranged in chronological order. A real-time intelligent processing module is constructed at the edge, and a task set is set for radar data analysis, traffic event identification and data uploading. An end-to-end latency threshold is set for each task using a multi-level inference scheduling unit, and a traffic complexity index is calculated through an adaptive computing power allocation mechanism to obtain the scheduling strategy. An improved BiSeNet network was constructed, which spatially weighted the multi-source perception data sequence by radar-guided spatial attention, introduced a multi-event target decoupling branch for traffic feature extraction, and performed cross-modal fusion in the radar-visual dynamic fusion head to obtain traffic target recognition results and traffic event feature maps. Based on the scheduling strategy, traffic target trajectories, lane occupancy information, and traffic event judgment results are generated according to the traffic target identification results and traffic event feature maps, forming structured traffic target data and event data; Structured traffic target data and event data, along with corresponding key video frames, are stored on local media and uploaded to the central platform via a communication module.

3. The method for real-time traffic target recognition based on edge computing in a radar-visual integrated machine according to claim 2, characterized in that, The initialization equipment includes a millimeter-wave radar for the radar vision integrated machine, a traffic monitoring camera, and a communication module.

4. The method for real-time traffic target recognition based on edge computing in a radar-visual integrated machine according to claim 2, characterized in that, The generation of the radar guidance map and preprocessed video frames includes: In the edge computing terminal, the millimeter-wave radar, traffic monitoring camera and communication module are powered on and started. The working mode, transmission power, sampling frequency and clock synchronization parameters of the millimeter-wave radar and the resolution, exposure time, frame rate and clock synchronization parameters of the traffic monitoring camera are initialized and configured so that the millimeter-wave radar and traffic monitoring camera work based on a unified time reference. The millimeter-wave radar is controlled to collect raw radar echo data covering the road monitoring area according to the scanning cycle, and encapsulated into continuous radar frame data according to a unified timestamp. The traffic monitoring camera is controlled to collect traffic scene video images according to the frame rate, and the collected image frames are marked according to the timestamp corresponding to the radar frame data to form a video frame sequence that is aligned with the radar frame data in time. The radar frame data is filtered, corrected, and target point cloud is extracted. Based on the calibration parameters stored in the edge computing terminal, the three-dimensional coordinates of the target points are mapped to image plane coordinates to generate a radar guidance map that is spatially aligned with the video frame sequence. The video frame sequence is then subjected to size scaling, color space conversion, and pixel value normalization to generate preprocessed video frames corresponding to the radar guidance map.

5. The method for real-time traffic target recognition based on edge computing in a radar-visual integrated machine according to claim 2, characterized in that, The formation of the multi-source sensing data sequence arranged in chronological order includes: The configured road area contour parameters and lane model parameters are read in the edge computing terminal. The lane model parameters include the lane center line, lane boundary line and lane number information. The road area contour parameters and lane model parameters are loaded into the scene of the region of interest division. The radar guidance map and pre-processed video frames are mapped to the coordinate system corresponding to the scene. The road monitoring area is extracted based on the road area contour parameters. Based on the lane model parameters, the corresponding lane interest area and roadside interest area are divided on the radar guidance map and pre-processed video frames, and area identifiers are assigned to each interest area. According to the acquisition time sequence, the radar guidance map after region of interest division and the pre-processed video frame are paired. Data with the same timestamp and region identifier are combined into a frame of multi-source sensing data. Multiple consecutive frames of multi-source sensing data are stored in the buffer in time order to form a multi-source sensing data sequence arranged in time order.

6. The method for real-time traffic target recognition based on edge computing in a radar-visual integrated machine according to claim 2, characterized in that, The obtained scheduling strategy includes: In the edge-side real-time intelligent processing module, a task set is created, which includes radar data analysis tasks, traffic event recognition tasks, and data upload tasks. Input data type and output data type are defined for each task. Each task is assigned an end-to-end latency threshold by a multi-level inference scheduling unit. The end-to-end latency threshold includes the maximum allowable processing time from the entry of multi-source sensing data into the edge computing terminal to the generation of the task output result. The task is divided into a multi-level priority queue according to the real-time requirements of different tasks, and a multi-level inference scheduling table is generated. Through an adaptive computing power allocation mechanism, the number of traffic targets, lane occupancy ratio, and environmental state parameters within the current time window are statistically analyzed from the multi-source perception data sequence. The traffic complexity index is calculated by combining the number of traffic targets, lane occupancy ratio, and environmental state parameters according to preset weight coefficients. The traffic complexity index is equal to the number of traffic targets multiplied by the first weight coefficient, plus the lane occupancy ratio multiplied by the second weight coefficient, plus the environmental state parameters multiplied by the third weight coefficient. Based on the traffic complexity index and the multi-level inference scheduling table, the inference frame rate, processing time slice length, and allocated computing power of radar data analysis tasks, traffic event identification tasks, and data upload tasks are dynamically adjusted to generate a scheduling strategy that includes task execution order, task execution cycle, and resource allocation ratio.

7. A method for real-time traffic target recognition based on edge computing using a radar-visual integrated machine according to claim 2, characterized in that, The process of obtaining traffic target recognition results and traffic event feature maps includes: Preprocessed video frames are read from multi-source sensing data sequences and input into the spatial and contextual paths of the improved BiSeNet network to obtain video feature maps at different scales. At the same time, the radar guidance map at the corresponding time is aligned to the size of the video feature map in terms of spatial scale. Radar guidance spatial attention calculation is performed on the radar guidance map. The attention weight coefficient of each spatial location is calculated, and the attention weight coefficient is multiplied with the video feature map at the corresponding spatial location to generate a video feature map that has undergone radar guidance spatial weighting. The video feature map, which has undergone radar-guided spatial weighting, is fused and compressed to generate a visual semantic feature map. The radar-guided map is then subjected to convolutional transformation and scale alignment to generate a radar semantic feature map. The visual semantic feature map and the radar semantic feature map are input into the radar-visual dynamic fusion head. The fusion weight is calculated based on the radar semantic feature map, and the visual semantic feature map is weighted to obtain a cross-modal fusion feature map. The cross-modal fusion feature map is input into the multi-event target decoupling branch, and the target basic feature map is output through convolutional layers and classification layers. The first decoupling sub-branch outputs the traffic target category feature map based on the target basic feature map. The second decoupling sub-branch uses the traffic target category feature map as a query to perform attention filtering on the target basic feature map and outputs the traffic target behavior state feature map. The third decoupling branch uses the traffic target category feature map and the traffic target behavior state feature map as context to perform risk attention modeling on the target basic feature map and outputs the traffic event risk feature map. By combining traffic target category feature maps, traffic target behavior state feature maps, and traffic event risk feature maps with cross-modal fusion feature maps, pixel-level or region-level decisions are made to generate traffic target recognition results and traffic event feature maps.

8. A method for real-time traffic target recognition based on edge computing using a radar-visual integrated machine, as described in claim 2, is characterized in that... The formation of structured traffic target data and event data includes: Pixel-level target masks are extracted from traffic target recognition results and traffic event feature maps. Pixel sets with the same target category label and spatial connectivity are clustered into traffic target instances. A frame-level target list containing target number, target category, pixel region and timestamp is generated. Range, speed and azimuth information are extracted from radar data with the same timestamp. The frame-level target list is matched with radar range, velocity, and azimuth according to timestamp and spatial location. Range, velocity, and azimuth values ​​are added to each traffic target instance. Target instances with similar spatial locations and the same target category in adjacent frames are associated in chronological order to generate continuous target trajectories. The continuous target trajectory includes target number, trajectory coordinate sequence, velocity sequence, and time sequence. Based on the lane model parameters, the trajectory coordinates are mapped to the lane coordinate system. The occupancy time of each lane is accumulated within the statistical time, and the lane occupancy ratio is calculated. The lane occupancy ratio is equal to the cumulative lane occupancy time divided by the statistical time length. Based on the comparison of speed sequence with speed limit threshold, comparison of travel direction with lane direction, and trajectory dwell time, speeding, wrong-way driving, long-term lane occupation, and littering events are determined. The target trajectory, lane occupancy ratio, and event records are encapsulated into structured traffic target data and structured traffic event data.

9. A method for real-time traffic target recognition based on edge computing using a radar-visual integrated machine according to claim 2, characterized in that, The process of storing structured traffic target data and event data, along with corresponding key video frames, on a local medium and uploading them to the central platform via a communication module includes: In the edge computing terminal, a unique index identifier is assigned to the structured traffic target data, structured traffic event data and key video frames. The index identifier is combined with the timestamp, device number and lane number to generate a storage file name or record number, and the data is written to the cache area in the local storage medium. The data in the buffer is stored and managed in a cyclical manner according to time order. When the space occupied by the buffer exceeds the limit, the earliest data record is deleted, and only the structured traffic target data, structured traffic event data and key video frames in the most recent time window are retained. Structured traffic target data, structured traffic event data, and key video frames are read from the cache according to the upload cycle and priority order, packaged into upload data frames, and sent to the central platform via wireless link.