Traffic event illegal behavior early warning system based on ai and multi-source data fusion

By distinguishing between duplicate and complementary video information in a multi-camera monitoring system and combining multi-source traffic data for fusion analysis, and dynamically adjusting similarity thresholds and local partition corrections, the problems of missed detections and false alarms in multi-camera monitoring are solved, achieving more efficient detection and early warning of traffic violations.

CN120032522BActive Publication Date: 2025-10-21JIANGSU SOUTHEAST INTELLIGENT TECH GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510487848.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-10-21
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify traffic violations in multi-camera monitoring scenarios, exhibiting high rates of missed detections and false alarms. In particular, when multiple cameras overlap or partially obstruct the coverage area, the lack of effective difference analysis and adjacent zone collaborative correction mechanisms negatively impacts traffic monitoring efficiency and safety management effectiveness.

Method used

By acquiring video data from multiple cameras and external traffic data, timestamp alignment and image feature extraction are performed to distinguish between duplicate and complementary video information. Combined with multi-source traffic data for fusion analysis, traffic violation indication information is generated. The similarity threshold is dynamically adjusted, and a local partitioning environmental complexity index and collaborative correction mechanism are adopted to improve detection accuracy.

Benefits of technology

It enhances the vehicle's continuous tracking capability across cameras, reduces missed detections and false judgments, and can flexibly determine the similarity and difference of images in complex environments, achieving more refined multi-camera fusion monitoring and improving the detection accuracy and timely warning of traffic violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032522B_ABST
    Figure CN120032522B_ABST
Patent Text Reader

Abstract

The application provides a traffic event illegal behavior early warning system based on AI and multi-source data fusion, relates to the technical field of intelligent traffic management, realizes timestamp alignment, image feature extraction and difference analysis by acquiring multiple camera videos and external traffic data, distinguishes repeated video information and complementary video information, and automatically and dynamically adjusts a similarity threshold according to traffic flow, weather, vehicle speed and the like; spatial correction between cameras is carried out by using an automatic calibration subunit, multi-part management is carried out on a target road on the basis of a detection result, and cooperative correction is carried out on an adjacent partition overlapping area, so that the robustness of cross-camera tracking and illegal behavior identification is improved; the system can also combine road geometric information to compare vehicle trajectories, accurately detect behaviors such as overspeeding, illegal lane changing and reverse driving, and output corresponding illegal early warning signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent traffic management technology, and more specifically, to a traffic incident violation warning system based on AI and multi-source data fusion. Background Art

[0002] In the field of intelligent traffic management, it is usually necessary to continuously monitor vehicles and traffic incidents within the target road area in order to promptly detect and warn of traffic violations or potential safety risks. In the existing technology, a single camera or a combination of multiple cameras is often used for traffic video acquisition. Although it can achieve a certain degree of illegal capture and traffic flow analysis, there are still problems such as inaccurate image similarity judgment, inability to dynamically adapt to different traffic flows and weather environments, and data redundancy in overlapping areas between cameras or insufficient blind spot identification, resulting in high missed detection and false alarm rates. In particular, when there is overlap or partial occlusion of multiple cameras in the coverage area, there is a lack of effective difference analysis and adjacent partition collaborative correction mechanism, making it difficult to accurately obtain the complete trajectory of the vehicle and identify potential violations in real time, affecting the overall traffic monitoring efficiency and safety management effect. Therefore, there is a need for a solution that can adaptively integrate external traffic data in a multi-camera monitoring scenario to further improve the detection accuracy of traffic incident violations and the timeliness of warnings. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention provides a traffic violation warning system based on AI and multi-source data fusion, including:

[0004] an acquisition unit, which acquires video data collected by multiple cameras deployed in a target road area and multi-source traffic data provided by at least one external data source;

[0005] A detection unit, based on the video data, determines camera images with overlapping coverage areas and generates a detection result; wherein the detection result includes repeated video information and complementary video information;

[0006] a fusion unit, which performs fusion analysis on the detection result and the multi-source traffic data to generate traffic violation indication information;

[0007] The early warning unit generates an early warning signal based on the traffic violation indication information; wherein the early warning signal is used to indicate potential and / or ongoing traffic violation events.

[0008] As an optional implementation, generating the detection result includes:

[0009] Performing time stamp alignment on the video data collected by the multiple cameras to generate multiple image frame sets;

[0010] For the image frame set, extract image features using a preset model to generate a feature vector for evaluating picture similarity;

[0011] Determining, based on the feature vector, image similarities between multiple image frames in the same image frame set;

[0012] In response to the picture similarity score being greater than or equal to a first threshold, marking the corresponding image frame as duplicate video information;

[0013] performing difference analysis on the image frames whose picture similarity scores are greater than or equal to the second threshold and less than the first threshold, and generating complementary video information;

[0014] The image frames marked as repeated video information and complementary video information are integrated to generate the detection result, and the detection result is output to the fusion unit.

[0015] As an optional implementation, the detection unit further includes:

[0016] The automatic calibration subunit obtains the spatial mapping relationship of the multiple cameras based on the feature points of the target road area, and performs spatial correction on the video data before calculating the picture similarity.

[0017] As an optional implementation manner, the detection unit performs a first dynamic adjustment on the first threshold and the second threshold based on the multi-source traffic data;

[0018] The multi-source traffic data includes: traffic volume, weather conditions and vehicle speed;

[0019] Among them, the traffic flow is obtained through radar detection data, the weather conditions are obtained through meteorological data, and the vehicle speed is obtained through vehicle network information.

[0020] As an optional implementation manner, the detection unit is further configured to:

[0021] Based on the detection result, the target road area is divided into a plurality of camera partitions; wherein the cameras corresponding to the image frames whose picture similarity scores are greater than or equal to a second threshold are considered to be in the same partition;

[0022] In each of the camera zones, calculating a zone environment complexity index of the camera zone based on local multi-source traffic data;

[0023] Based on the partition environment complexity index, performing a second dynamic adjustment on the first threshold and the second threshold;

[0024] The first threshold and the second threshold of different camera partitions are independent of each other.

[0025] As an optional implementation manner, the detection unit is further configured to:

[0026] When overlapping areas are detected between different partitions, the corresponding partitions are marked as adjacent partitions;

[0027] For the adjacent partitions, the similarity determination and difference analysis results are collaboratively corrected.

[0028] As an optional implementation, the collaborative correction includes:

[0029] Based on the degree of overlap of coverage areas between the adjacent partitions, performing confidence weighting processing on the similarity determination and the difference analysis results;

[0030] In response to detecting that the same target vehicle and / or the same frame image generates a conflict determination in the adjacent partitions, recalculating the similarity score and / or the difference score through cross-validation;

[0031] In response to the completion of the recalculation, the repeated video information and the complementary video information are marked as updated, and the updated detection result is output.

[0032] As an optional implementation, the multi-source traffic data further includes: road geometry information; and the fusing and analyzing the detection result with the multi-source traffic data to generate traffic violation indication information includes:

[0033] Based on the image frames marked as repeated video information, the consistency of the target vehicle's continuous tracking trajectory in the overlapping area covered by multiple cameras is confirmed, and the repeated video information is used as the vehicle trajectory benchmark;

[0034] Based on the image frames marked with complementary video information, the moving position of the target vehicle under different viewing angles or partial occlusion conditions is identified to generate vehicle trajectory data;

[0035] Aligning and comparing the vehicle trajectory data with road geometry information to generate a comparison result;

[0036] Based on the comparison result, determining whether the target vehicle has committed a preset illegal behavior;

[0037] In response to the target vehicle having the preset illegal behavior, corresponding vehicle identification information and timestamp are extracted, and indication information for indicating the illegal behavior is generated.

[0038] As an optional implementation manner, generating complementary video information includes:

[0039] Performing foreground extraction processing on the image frame to determine a suspicious target area;

[0040] Matching the suspicious target area with the vehicle position and / or lane information in the multi-source traffic data to generate a difference parameter;

[0041] In response to the difference parameter being greater than or equal to a preset determination parameter, marking the corresponding image frame as complementary video information;

[0042] In response to the difference parameter being smaller than the preset determination parameter, the corresponding image frame is updated to repeated video information.

[0043] Compared with the existing technology, this application improves the ability to continuously track vehicles across cameras by introducing the distinction and application of repeated video information and complementary video information in a multi-camera networking scenario, reduces missed detections caused by repeated identification or occlusion of the same target, and can obtain a more complete target trajectory through the complementarity of different perspectives. The mechanism of dynamically adjusting the similarity threshold based on multi-source traffic data enables the system to flexibly determine the similarity and difference of the images in complex environments such as a sudden increase in traffic volume, bad weather, or high-speed driving of vehicles, thereby reducing the misjudgment rate caused by environmental factors. On this basis, the environmental complexity index of the local partition is used to independently adjust the threshold, and a collaborative correction method is used in the overlapping areas between partitions, so as to achieve a more refined multi-camera fusion monitoring strategy, which not only ensures robustness on busy roads, but also avoids waste of resources and unnecessary false alarms in areas with low traffic volume. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A schematic diagram of a traffic incident violation warning system based on AI and multi-source data fusion provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of camera partitioning provided in an embodiment of the present application;

[0046] Figure 3 A flowchart of a method for generating complementary video information provided in an embodiment of the present application.

[0047] Reference numerals: 10, acquisition unit; 20, detection unit; 30, fusion unit; 40, early warning unit. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0049] See also Figure 1FIG. 1 is a schematic diagram of a traffic incident violation warning system based on AI and multi-source data fusion according to an embodiment of the present application. The system includes an acquisition unit 10, a detection unit 20, a fusion unit 30, and a warning unit 40, wherein:

[0050] An acquisition unit 10 acquires video data collected by multiple cameras deployed in a target road area and multi-source traffic data provided by at least one external data source;

[0051] The detection unit 20 determines, based on the video data, camera images with overlapping coverage areas and generates a detection result; wherein the detection result includes duplicate video information and complementary video information;

[0052] A fusion unit 30 performs fusion analysis on the detection result and the multi-source traffic data to generate traffic violation indication information;

[0053] The warning unit 40 generates a warning signal based on the traffic violation indication information; wherein the warning signal is used to indicate potential and / or ongoing traffic violation events.

[0054] This application can be deployed in a traffic monitoring center or cloud server to quickly detect and warn potential or ongoing traffic violations in a target road area. Multi-source traffic data may include but is not limited to radar detection data, meteorological data, vehicle network information, map data, etc.

[0055] When roads are large or complex, relying solely on a single camera (regardless of its high resolution) still struggles to fully capture the continuous movement of vehicles at different locations. This can easily lead to missed detections and false detections when vehicles are partially obscured or when visibility is reduced at night. Furthermore, in multi-camera networking scenarios, lacking overlap analysis of camera coverage areas often results in the same target vehicle being independently identified multiple times, preventing the system from making a comprehensive assessment across cameras. This results in inefficient and inaccurate identification and evidence collection.

[0056] This application emphasizes the use of the detection unit 20 to perform overlapping coverage area analysis on the video data collected by multiple cameras, distinguishing between repeated video information and complementary video information, thereby strengthening the continuous tracking of the same vehicle (reducing repeated calculations and redundant analysis), and using the complementary information provided by different perspectives to compensate for blocked areas and blind spots, thereby improving the accuracy of detection of illegal acts.

[0057] In addition, general video analysis algorithms judge the behavior of vehicles or pedestrians on the road, and often lack effective integration of external environmental or traffic factors such as weather, traffic flow, and vehicle speed.

[0058] For example, in heavy rain, heavy snow, and sandstorms, the camera's image clarity and lighting conditions will be seriously affected; in conditions of abnormally dense traffic, single video analysis is prone to decreased detection performance or increased misjudgment rate.

[0059] This application integrates external data sources (such as radar detection data, meteorological data, and vehicle network information) with video detection results to provide a more comprehensive and dynamic assessment of traffic conditions. For example, in the event of a sudden surge in traffic volume or deteriorating weather, the similarity threshold can be appropriately raised or lowered, and the detection area can be expanded or reduced. Furthermore, when determining violations, multi-dimensional data such as vehicle speed and lane departure can be incorporated to reduce missed detections or misjudgments caused by visual interference.

[0060] For the above acquisition unit 10:

[0061] In practice, multiple high-definition cameras (optionally with infrared or night vision capabilities) can be installed at key arteries and intersections in the target road area to provide 24-hour monitoring. The cameras can be connected to the Internet via fiber optics, 5G networks, or wired Ethernet to transmit the surveillance video stream in real time to a data center or cloud.

[0062] At the external data source level, it can connect to the traffic management department's Internet of Vehicles platform, road radar detection system, meteorological department's weather database, etc., and obtain corresponding multi-source data through API interface or data subscription.

[0063] The acquisition unit 10 can be implemented by a high-performance server or edge computing device, which is responsible for performing preliminary timestamp or location index association on the above-mentioned video stream and multi-source traffic data to ensure that synchronous or approximately synchronous processing can be achieved in the subsequent detection and analysis link.

[0064] Regarding the above detection unit 20:

[0065] In a specific implementation, the detection unit 20 can use a computer vision algorithm on a server or cloud to identify possible overlapping areas in different camera images. By comparing image content or road geographic location information, it can determine which cameras' monitoring ranges overlap in space.

[0066] To improve accuracy, the detection unit 20 can combine camera calibration information (such as installation altitude, installation angle, and lens focal length) with map information of the target road area to perform similarity analysis on the frames using image recognition algorithms and feature matching algorithms (such as SIFT, SURF, and ORB). When two or more video frames are detected showing similar content within the same field of view (e.g., the same lane or the same vehicle), these frames are identified as duplicate video information for subsequent use in confirming vehicle trajectory consistency or locating violations.

[0067] On the other hand, images that can complement each other's target vehicle information from different perspectives are identified as complementary video information to help the algorithm obtain complete vehicle trajectory information even in partial occlusion, insufficient light, and other conditions.

[0068] Regarding the above fusion unit 30:

[0069] In a specific implementation, the fusion unit 30 can make a comprehensive judgment based on the detection results from different cameras and the traffic flow, weather, vehicle speed and other information provided by external data sources based on AI models (such as deep learning networks, expert systems based on fusion rules, etc.).

[0070] Vehicle cross-camera matching is performed on repeated video information to confirm whether the trajectory of the same vehicle in different images is continuous and whether there are any illegal behaviors such as abnormal lane changes, speeding or crossing the line.

[0071] Complementary video information is analyzed to identify violations that may be obscured or unrecognized in a single camera's field of view. Traffic flow and connected vehicle GPS information from multiple data sources are then matched with vehicle identification information in the video to provide a more comprehensive picture of road traffic conditions. This multi-data fusion approach significantly improves the accuracy of violation identification and reduces the rate of false positives.

[0072] Regarding the above-mentioned early warning unit 40:

[0073] In practice, when the fusion unit 30 identifies a possible traffic violation (such as illegal lane change, speeding, running a red light, etc.), the warning unit 40 generates a corresponding warning signal. This warning signal can be transmitted via the traffic control center's large monitoring screen, mobile app, SMS alert, and other means.

[0074] In terms of potential violation warnings, the system can flag vehicles that haven't clearly violated regulations but are exhibiting high-risk driving behavior in real time. For example, in rainy or snowy weather, at high speeds, and with heavy traffic, if a vehicle is detected changing lanes frequently or driving too close to the vehicle ahead, a warning message can be sent to the backend or roadside display screens.

[0075] In terms of detecting violations that have already occurred, the system can automatically capture or extract relevant video clips of violations, and upload them together with vehicle license plate, timestamp, location and other information to the database of the relevant department for subsequent punishment or intervention.

[0076] It should be noted that the functions of each unit can be performed by different functional modules on a single server or cloud platform, or preliminary video analysis can be performed on the edge side and the results can be uploaded to the cloud for secondary fusion.

[0077] Among them, the acquisition unit 10 needs to be connected to the SDK or network flow protocol of the camera manufacturer; the acquisition of external data sources can achieve real-time or quasi-real-time data interaction based on RESTful API, MQTT message queue or WebSocket.

[0078] The core of the detection unit 20 lies in computer vision and image processing algorithms, for example, it may include a feature extraction model based on convolutional neural network (CNN), a picture similarity calculation module based on traditional feature point matching, and automatic judgment logic for overlapping areas.

[0079] The fusion unit 30 can be based on a big data platform or an AI analysis platform, and can process spatiotemporal data from multiple sources, and conduct a comprehensive comparison of vehicle trajectories, vehicle speeds, and lane usage, and ultimately derive information indicating illegal behavior.

[0080] The early warning unit 40 can push abnormal information or potential risk information to the traffic control management terminal or other user terminals in the first time based on event-driven message services.

[0081] Thus, the system provided by this application can achieve full-time, all-around monitoring and early warning of road traffic violations based on the functional collaboration of the aforementioned units. Through multi-source data fusion and intelligent analysis by AI algorithms, the system has advantages in improving traffic management efficiency, reducing the burden of manual inspections, and lowering the risk of traffic accidents. The system is particularly suitable for urban roads, highways, and other traffic-intensive scenarios.

[0082] As an optional implementation, generating the detection result includes:

[0083] Performing time stamp alignment on the video data collected by the multiple cameras to generate multiple image frame sets;

[0084] For the image frame set, extract image features using a preset model to generate a feature vector for evaluating picture similarity;

[0085] Determining, based on the feature vector, image similarities between multiple image frames in the same image frame set;

[0086] In response to the picture similarity score being greater than or equal to a first threshold, marking the corresponding image frame as duplicate video information;

[0087] performing difference analysis on the image frames whose picture similarity scores are greater than or equal to the second threshold and less than the first threshold, and generating complementary video information;

[0088] The image frames marked as repeated video information and complementary video information are integrated to generate the detection result, and the detection result is output to the fusion unit 30.

[0089] In practice, the system first needs to timestamp-align the raw video streams captured by multiple cameras to form a comparable set of image frames. To ensure time synchronization between image frames across different cameras, the Network Time Protocol (NTP) or other distributed synchronization mechanisms can be used during camera deployment or on the server side to uniformly calibrate the video streams of each camera.

[0090] Image frames acquired at different times are grouped according to their corresponding timestamps. Image frames with adjacent timestamps (e.g., at the same capture time or in the same batch) can be merged into a single image frame set.

[0091] Furthermore, in some cases, alignment error correction can be performed (for example, through interpolation or discarding invalid or redundant frames) to ensure that the physical scenes corresponding to the images from each camera have consistent or nearly consistent time windows during subsequent similarity analysis. This timestamp alignment process allows image frames from different cameras at the same moment or time window to be packaged into several image frame sets, providing the foundational data structure for subsequent feature extraction and similarity analysis.

[0092] After obtaining the above-mentioned image frame set, the system will call the pre-configured image feature extraction model to process these frames and generate feature vectors for evaluating the similarity of the pictures.

[0093] When choosing a preset model, you can extract key features of the image based on traditional feature point matching algorithms (such as SIFT, SURF, ORB) or convolutional neural network (CNN) models based on deep learning.

[0094] For each image frame, appearance features (such as color histogram, edges, and texture) and / or semantic features (such as road position, vehicle bounding box, and license plate information) are extracted and packaged into a feature vector for output to subsequent modules. After generating the feature vector, the detection unit 20 performs pairwise or multi-pair comparisons on multiple frames in the same image frame set (from different cameras) to obtain a similarity score.

[0095] Formulas such as cosine similarity, Euclidean distance, and Mahalanobis distance can be used to measure the similarity between feature vectors. Metric learning within deep learning models can also be used to measure similarity. The calculated similarity for each pair (or group) of frames is then output as a numerical result, such as a score between 0 and 1, or presented as a percentage.

[0096] In a specific implementation, once the image similarity score between certain image frames is detected to be greater than or equal to a first threshold, it indicates that the image group is highly likely to belong to the same road section, the same vehicle (or the same target), and the same viewpoint or scene. The system directly marks the corresponding image frames as duplicate video information. Furthermore, this information can be stored in a database, along with the corresponding timestamp, camera ID, and other information, to facilitate subsequent cross-camera tracking and consistency verification by the fusion unit 30.

[0097] For image frames with similarity scores between the second and first thresholds, it indicates that they do not highly overlap with the reference frame, but still have many common features (perhaps the same vehicle from a different perspective or a similar scene from a partially occluded perspective). At this point, the system will further perform a difference analysis;

[0098] In specific implementations, dissimilarity metrics can be calculated based on information such as target outline, deformation, vehicle features (e.g., license plate, roof shape), and local texture. Vehicle speed and position from external data sources can also be referenced and matched against the frame information to confirm whether they represent the same vehicle / object. When dissimilarity analysis confirms that the frames complement each other in terms of perspective, vehicle position, and other aspects, the system marks them as complementary video information, allowing for supplementary or correction of vehicle trajectories in subsequent fusion analysis.

[0099] After completing the above-mentioned marking operation, the detection unit 20 integrates the image frames corresponding to the repeated video information and the complementary video information to form a final detection result. The detection result includes:

[0100] A list of frame pairs with high similarity (duplicate video information) to facilitate subsequent cross-camera consistent vehicle recognition;

[0101] The frame pairs or frame groups (complementary video information) obtained after difference analysis can make up for the monitoring blind spots or occlusion problems under a single camera.

[0102] For example, the system may create a structured data record including the corresponding image frame ID, similarity score, difference analysis parameters, and a label indicating whether it is marked as "duplicate" or "complementary".

[0103] Finally, the detection unit 20 sends the above-mentioned integrated detection results to the fusion unit 30 to combine more external data for in-depth illegal behavior identification.

[0104] The output may be in the form of a database interface, a message queue or a Web API, to ensure that the fusion unit 30 can receive the detection results in the first place and perform subsequent cross-camera vehicle tracking, violation determination and early warning processing.

[0105] In actual deployment, if there are low bandwidth or network delay issues, local edge computing nodes can be used to quickly determine the similarity and difference directly, and then only send the final marking information to the cloud, reducing data transmission and improving processing efficiency.

[0106] By setting first and second thresholds and employing a difference analysis step, it is possible to effectively distinguish between highly similar frames (truly duplicate information) and frames with a certain degree of similarity from different perspectives but complementary information (complementary video information), thereby improving overall detection accuracy. For frames between the second and first thresholds, difference analysis can further determine whether there are issues such as partial occlusion or angular offset, rather than simply ignoring such frames and thus missing potential violation clues. The duplicate and complementary video information are then integrated and output to the fusion unit 30, facilitating subsequent cross-camera vehicle identification, trajectory splicing, and integration with other multi-source data for higher-level violation determination and early warning.

[0107] As an optional embodiment, the detection unit 20 also includes an automatic calibration subunit, which can obtain the spatial mapping relationship between multiple cameras based on feature points in the target road area and perform necessary spatial correction on the video data before calculating the picture similarity.

[0108] In practice, several calibration feature points are first placed in the target road area, such as road markings, traffic signs, curb corners, or specially placed markers. By mapping these feature points or combining them with drone photography, lidar scanning, and other methods, their precise position information in the world coordinate system can be obtained.

[0109] After receiving images from different cameras, the automatic calibration subunit uses computer vision algorithms to identify the aforementioned feature points in the images and match them with real-world coordinates. It then calculates the extrinsic and intrinsic parameters of each camera using multi-view geometry or camera calibration algorithms (such as the Zhang Zhengyou calibration method). Based on this calibration result, the automatic calibration subunit performs projection transformation and distortion correction on the images captured by each camera, so that the images of the same vehicle or road section captured by different cameras can be mapped to the same reference coordinate system or top-down plane.

[0110] In this way, when performing similarity analysis, the differences caused by camera installation position, shooting angle or lens distortion can be eliminated, thereby improving the accuracy of recognizing overlapping images in the coverage area.

[0111] The present application further proposes another optional implementation method, namely, in the detection unit 20, the system is able to perform a first dynamic adjustment of the first threshold and the second threshold based on multi-source traffic data. The so-called first threshold and the second threshold are used to distinguish whether the similarity score between image frames is sufficient to mark them as duplicate video information or complementary video information, respectively. However, in actual environments, factors such as traffic flow, weather conditions, and vehicle speed often affect the quality and difficulty of image recognition. To this end, when the detection unit 20 receives real-time traffic flow, weather, and vehicle network information from external data sources, it can automatically modify the threshold.

[0112] For example, when traffic volume increases sharply, the density of vehicles in the picture increases and occlusions become more frequent. The system can appropriately lower the similarity threshold, so that there is a more tolerant judgment range for frames that may belong to the same scene. In severe weather conditions, such as heavy rain, dense fog or dust, the picture quality is significantly impaired, and the threshold can be further lowered or raised in combination with meteorological data to reduce misjudgment caused by noise interference.

[0113] In addition, the real-time driving speed of vehicles in the Internet of Vehicles information also affects the judgment of picture similarity. If it is detected that vehicles are generally passing at high speed, the vehicle position in the picture frame changes faster. The system can also adjust the threshold in time to avoid missing frames with shorter time domain overlap.

[0114] In this way, adaptive adjustment of the similarity threshold is achieved in complex and changeable actual road environments, thereby taking into account both accuracy and real-time performance and improving the robustness of multi-camera fusion monitoring.

[0115] It is understandable that the configuration of the above-mentioned automatic calibration sub-unit enables the images of different cameras to obtain unified spatial correction before entering the similarity analysis process, fundamentally reducing the errors caused by installation angles and lens differences; and the mechanism of dynamic adjustment of thresholds based on multi-source traffic data maintains the rationality and adaptability of similarity division when dealing with bad weather, sudden increases in traffic volume or abnormal vehicle speeds.

[0116] In an optional embodiment, the detection unit 20 first partitions the target road area based on the aforementioned detection results, which include duplicate video information and complementary video information obtained by analyzing the similarity or difference between different camera images based on the first threshold and the second threshold.

[0117] At this time, if the picture similarity score between certain image frames is greater than or equal to the second threshold, it means that there is a certain degree of overlap or complementary area in the monitoring range between the corresponding cameras. These cameras can be regarded as the same partition and classified into the same camera partition.

[0118] In other words, the detection unit 20 utilizes the camera IDs associated with these high-similarity or relatively high-similarity frames to aggregate cameras with similar or obvious overlapping monitoring ranges into a partition, thereby achieving a preliminary spatial division of the target road area.

[0119] After completing the partitioning, the detection unit 20 calculates the partition environment complexity index within each camera partition based on the local multi-source traffic data.

[0120] Compared with the data of the entire area, local multi-source traffic data usually refers to the road conditions actually monitored or counted within the camera partition, including local traffic volume, local weather conditions (such as local heavy rain or dense fog sections), average vehicle speed and congestion level within the partition.

[0121] For example, if traffic volume in a zone is high and weather is inclement, the camera image may appear densely populated with vehicles and exhibit significant noise, posing challenges to both similarity analysis and object recognition. In this case, the zone's environmental complexity index will be higher. On the other hand, if traffic volume in the zone is low, the weather is fine, and vehicle movement is stable, the complexity index will be lower.

[0122] After calculating the environment complexity index of each partition, the detection unit 20 may perform a second dynamic adjustment on the first threshold and the second threshold according to the index.

[0123] It should be noted that, unlike the first dynamic adjustment based on global multi-source traffic data, the second dynamic adjustment focuses more on the real-time status of local partitions. Each partition can independently modify the judgment threshold originally used to divide duplicate video information and complementary video information according to its environmental complexity index.

[0124] Specifically, if the environmental complexity index of a certain partition is high, it means that it is necessary to be relatively more flexible in the judgment of image similarity to avoid being overly strict under high noise and high occlusion conditions, which may lead to missed detection or false detection; conversely, if the environmental complexity index of a certain partition is low, a stricter threshold can be set for the similarity judgment, thereby reducing redundant calculations or unnecessary misjudgments in the area.

[0125] It is important to note that the first and second thresholds for different camera zones are independent of each other, and each zone can have its own threshold configuration. This allows for better matching of the actual conditions of different road sections and different monitoring coverage areas.

[0126] For example, within the same city, some sections may experience chronically high traffic volume and be susceptible to inclement weather, while others may experience stable traffic flow and relatively mild weather. Forcibly applying a uniform threshold strategy could lead to misjudgments or oversensitivity in some sections. By subdividing the target road area and adjusting the threshold locally, the detection unit 20 can adapt to a wider range of road environments, significantly improving the accuracy and adaptability of the multi-camera fusion monitoring system.

[0127] In this way, by treating cameras with an image similarity score greater than or equal to the second threshold as belonging to the same zone, calculating the environmental complexity index for each zone, and dynamically adjusting the first and second thresholds accordingly, the detection unit 20 can implement a more flexible similarity determination mechanism at the local level. This mechanism not only enhances the system's robustness in scenarios such as high traffic volume, inclement weather, or complex terrain, but also provides a more accurate foundation for subsequent illegal behavior identification and real-time warnings.

[0128] For example, see Figure 2 , Figure 2 A schematic diagram of camera partitioning is provided in an embodiment of the present application. Several cameras are deployed for surveillance at the intersection of a main road and a branch road in a city. After automatically calibrating the installation positions, shooting angles, and coverage areas of the multiple cameras (or manually calibrating them in advance), the detection unit 20 can perform similarity calculations on the video frames captured by these cameras. For ease of illustration, this example involves five cameras, numbered C1, C2, C3, C4, and C5.

[0129] According to the above embodiment, the detection unit 20 performs similarity analysis on the image frame set within the same time window. Statistics show that:

[0130] The similarity scores between C1, C2, and C3 are higher than the second threshold at most moments;

[0131] There are also a large number of similarity scores between C4 and C5 that are greater than or equal to the second threshold, but the similarity scores between them and C1, C2, and C3 are generally low.

[0132] Because there is overlap or complementary information between C1, C2, and C3, and the similarity score is stably greater than the second threshold, the detection unit 20 divides C1, C2, and C3 into camera partition A;

[0133] For C4 and C5, although the similarity scores between them are high, the intersection with the camera coverage area in partition A (C1, C2, C3) is not large, and the similarity scores are not enough to meet the second threshold judgment standard. Therefore, C4 and C5 are divided into camera partition B.

[0134] As a result, the target road area is effectively divided into at least two relatively independent camera monitoring zones, namely zone A and zone B. The division result can be recorded in the system in a data structure or visual form, for example:

[0135] Partition A: contains cameras C1, C2, and C3;

[0136] Partition B: Contains cameras C4 and C5.

[0137] For each camera zone, the detection unit 20 combines local multi-source traffic data to calculate an Environmental Complexity Index (ECI). This index can range from 0 to 1 or 0 to 100, with higher values ​​indicating a more complex environment. The following are calculation examples for two zones:

[0138] Calculation of the environmental complexity index of partition A:

[0139] Regarding local traffic flow, within the area monitored by Zone A (e.g., a main road section), the traffic flow rate, as measured by roadside radar or microwave detectors, is 2,400 vehicles per hour (during peak hours). Compared to the local road design flow rate, this value exceeds the 80% threshold for saturation flow, indicating a high level of traffic complexity.

[0140] Regarding the weather conditions, the meteorological data report shows that it is currently raining moderately. Visibility is moderate, but waterlogging on the road will cause vehicles to slightly deviate from their driving trajectories, and raindrops and light reflections will interfere with the image.

[0141] As for the average vehicle speed, according to the statistics of the vehicle network information (V2X system), the average speed of vehicles in zone A is about 45 km / h, which is slightly lower than usual, indicating that the road has a certain congestion trend, but it has not yet been seriously congested.

[0142] The system performs weighted or rule-based operations on the above data and obtains the environmental complexity index ECI_A of partition A in the current period as 0.75 (when the value is between 0 and 1, 0.75 indicates medium to high complexity).

[0143] Calculation of the environmental complexity index of partition B:

[0144] As for local traffic flow, the branch roads covered by Zone B have relatively small traffic flow, only 800 vehicles per hour, and are relatively smooth.

[0145] Regarding the weather conditions, although it was also raining moderately, the drainage facilities on the branch roads were better, so the interference from accumulated water was not significant. The weather's impact on visibility was similar to that in Zone A.

[0146] As for the average vehicle speed, the Internet of Vehicles statistics show that the average speed of vehicles in zone B is about 60 km / h, and there is no serious congestion.

[0147] During the calculation, the system found that this partition was mainly affected by moderate rain and the traffic volume was not large. Therefore, after comprehensive weighting, the environmental complexity index ECI_B of partition B was 0.40 (the overall environmental interference was low).

[0148] Since the environmental complexity indexes of different partitions are significantly different, the detection unit 20 will make differential adjustments to the threshold strategies of the two partitions, thereby reducing missed detections or misjudgments in complex environments and avoiding wasting too many resources in simple scenarios.

[0149] For Partition A threshold adjustment:

[0150] Exemplarily, the initial first threshold T1_base and the second threshold T2_base (assuming T1_base=0.80, T2_base=0.60 respectively).

[0151] Combined with ECI_A=0.75, the system determines that zone A is in a highly congested and rainy environment. Vehicle overlap in the image may be significant and there is a lot of noise interference. To improve the ability to capture potential duplicate or complementary images, the threshold can be slightly lowered. For example, the updated first threshold T1_A=0.78; the updated second threshold T2_A=0.58;

[0152] In this way, partition A can more easily identify relatively similar image frames in the picture similarity score and regard them as repeated or complementary information, avoiding missed detections caused by traditional fixed thresholds.

[0153] For partition B threshold adjustment:

[0154] Also based on T1_base=0.80 and T2_base=0.60. T1_base is the first threshold of partition B, and T2_base is the second threshold of partition B;

[0155] Since ECI_B=0.40, the environmental interference is not significant, and the system can appropriately increase or maintain the threshold. For example, in partition B, the updated first threshold T1_B=0.82; the updated second threshold T2_B=0.61;

[0156] Compared with partition A, partition B has better image quality and greater vehicle dispersion, making it easier to distinguish between truly repeated and accidentally similar scenes. Therefore, using a slightly higher threshold can reduce misjudgments and improve efficiency.

[0157] This allows each partition to have a threshold setting tailored to its specific situation, enabling a locally refined image similarity determination strategy. Independent threshold adjustments between partitions avoid the drawbacks of a one-size-fits-all approach, which can easily lead to missed detections in highly complex scenes and over-capture redundant information in less complex ones.

[0158] This example demonstrates this technology on urban roads, but it can also be extended to highways, bridges, tunnels, or other scenarios with multi-camera coverage. As long as similarity analysis and partitioning are required, this method can be effectively applied.

[0159] It is understood that the specific calculation method for the aforementioned partitioned environment complexity index, the setting of the initial threshold, and the threshold update amplitude or adjustment rules for each partition can be achieved through various machine learning and big data analysis methods. For example, the system can use deep neural networks, random forests, or other AI models for training based on a large amount of historical monitoring data and real-world traffic records to automatically learn the optimal threshold adjustment strategy. It can also combine expert rules and heuristic algorithms to more specifically optimize the weight distribution in the complexity index calculation.

[0160] As an optional implementation manner, the detection unit 20 is further configured to:

[0161] When overlapping areas are detected between different partitions, the corresponding partitions are marked as adjacent partitions;

[0162] For the adjacent partitions, the similarity determination and difference analysis results are collaboratively corrected.

[0163] In actual road monitoring scenarios, although the detection unit 20 will divide the target road area into multiple camera partitions based on similarity scores and other results; however, in scenarios such as urban interchanges, complex road networks or multi-story elevated roads, different partitions are not always independent of each other, and some areas will often be monitored simultaneously under the coverage of two or more partitions.

[0164] To this end, the present application adds corresponding processing logic in the detection unit 20. When the system recognizes that there are overlapping areas between different partitions, these partitions will be marked as adjacent partitions. At the same time, in order to ensure the consistency of recognition across partition boundaries, the system will coordinate the similarity judgment and difference analysis results.

[0165] In specific implementations, the detection unit 20 first identifies subareas with persistently high spatial or temporal overlap based on camera coverage, image similarity information, and previously generated duplicate and complementary video information. If the cameras in two subareas repeatedly produce image frames with a similarity score greater than or equal to a predetermined threshold (e.g., a second threshold) at the same or adjacent times, the system determines that the monitored areas of these two subareas significantly overlap and marks them as adjacent. The system maintains an internal subarea relationship table, recording the adjacency of each subarea and the corresponding overlapping region coordinates or overlapping frame information.

[0166] Next, the detection unit 20 coordinates and calibrates the similarity determination and difference analysis results within these adjacent partitions. Specifically, each partition typically has independent threshold parameters and difference analysis mechanisms to adapt to the partition's own traffic flow, weather conditions, or camera configuration. However, when a vehicle or image spans adjacent partitions, if the two partitions have conflicting or inconsistent determinations of its similarity, occlusion status, or vehicle feature information (for example, video information marked as complementary in partition A but determined as duplicate in partition B), the system needs to conduct a unified review within the overlapping range of adjacent partitions, using multi-source data or other feature points for comparison and cross-validation.

[0167] During the collaborative correction process, the system determines which partitions have the most accurate recognition results by weighting confidence and recalculating similarity scores. If external data such as vehicle location, speed, and road geography confirm that two partitions are capturing the same vehicle or road section, the system will correct any previously conflicting similarity or difference results and, if necessary, update the markers for duplicate or complementary video information.

[0168] In this way, even if adjacent partitions use different threshold strategies or there are significant differences in shooting angles, collaborative correction can be used to maintain consistency in cross-partition recognition and reduce the risk of repeatedly identifying the same vehicle or misidentifying it as different objects.

[0169] For example, at the intersection of a city's elevated road and a ground auxiliary road, zone A mainly monitors the elevated lanes, while zone B is responsible for the ground auxiliary road. There is a certain overlap between the two in a certain section of the ramp area. At this time, a vehicle driving out at high speed will briefly appear in the picture of zone A, and then continue to enter the field of view of zone B. Since the installation angles and focal lengths of the two cameras are different, there may be some differences in the imaging of the vehicle's appearance, resulting in inconsistent classification or similarity scores of the vehicle within each zone. The system will trigger a collaborative correction process by identifying the overlapping areas here that belong to adjacent zones, and uniformly compare the frame sequences of the vehicle during the crossing of adjacent zones. If it is confirmed to be the same vehicle, zone A or zone B will adjust its recognition results accordingly and update the screen mark to form a more complete cross-zone trajectory for the vehicle.

[0170] In this way, the present application effectively avoids the problems of repeated identification, misjudgment or information fragmentation across partition boundaries, and further enhances the overall accuracy and coordination of multi-camera fusion monitoring. After marking adjacent partitions and performing collaborative correction, the detection unit 20 can output more reliable and consistent repeated video information and complementary video information, providing a more robust data basis for the subsequent fusion unit 30 and the early warning unit 40. Combined with the aforementioned technical means such as dynamic adjustment of thresholds, partition environment complexity index, and automatic calibration sub-units, the present application provides a more complete solution for the accurate detection and real-time early warning of various traffic violations in urban roads, elevated interchanges, and complex road network environments.

[0171] As an optional implementation, the collaborative correction includes:

[0172] Based on the degree of overlap of coverage areas between the adjacent partitions, performing confidence weighting processing on the similarity determination and the difference analysis results;

[0173] In response to detecting that the same target vehicle and / or the same frame image generates a conflict determination in the adjacent partitions, recalculating the similarity score and / or the difference score through cross-validation;

[0174] In response to the completion of the recalculation, the repeated video information and the complementary video information are marked as updated, and the updated detection result is output.

[0175] In specific implementation, within the processing framework of adjacent partitions, once it is confirmed that there is overlap in the monitoring areas between two or more camera partitions, it is necessary to perform more refined collaborative correction on the similarity determination and difference analysis results across partitions.

[0176] Specifically, the detection unit 20 first performs confidence weighting processing on the obtained similarity determination and difference analysis results based on the degree of overlap of coverage areas between adjacent partitions.

[0177] For example, if partition A and partition B have a large overlapping monitoring area and have maintained highly consistent judgment results in multiple time periods in the past, it means that both have strong monitoring capabilities for adjacent areas; at this time, they can be given high and similar confidence weights to ensure that no one side is favored during fusion.

[0178] On the contrary, if the camera clarity or coverage of a certain partition is relatively weak, its weight in the adjacent partition will be relatively reduced to avoid the overall result being affected by the misjudgment of a single path.

[0179] After completing the preliminary confidence weighting, if the system further detects that the same target vehicle or the same frame image has a conflict judgment between adjacent partitions, it will enter the cross-validation phase.

[0180] For example, partition A marks a vehicle as complementary video information, while partition B considers it to be duplicate video information; or the similarity score in partition A is low, while partition B gives a higher score, and there is a clear disagreement between the two.

[0181] At this time, the detection unit 20 can call external data or deeper visual features for cross-validation and recalculate the similarity score or difference score to eliminate deviations caused by factors such as partition threshold differences, different occlusion perspectives, or data synchronization errors.

[0182] For example, specific practices may include:

[0183] Perform secondary comparison on vehicle local features (license plate shape, body color / texture, feature point distribution);

[0184] Check vehicle network GPS information, vehicle radar marking data or road geometry information to confirm whether it is the same vehicle appearing at the same time and location;

[0185] According to the continuous motion trajectory of the previous and next frames, it is determined whether the recognition results of partition A and partition B are consistent in time and space.

[0186] After completing the cross-validation, the detection unit 20 will correct or update the duplicate video information and complementary video information labels previously generated in partition A and partition B based on the updated similarity score and / or difference score, and output the updated detection result.

[0187] The update process may include:

[0188] If it is confirmed that it is the same vehicle and should be marked as duplicate video information, the original mark of the complementary information will be deleted; conversely, if it is confirmed that it is a different vehicle or the perspectives are complementary, the complementary video information will be retained and the duplicate mark will be revoked.

[0189] To avoid the recurrence of the same or similar conflicts in the future, the system can record the cross-validation results in the database and adjust the confidence factor for partition A or partition B, making it easier to reach consensus in the same or similar situations in the future.

[0190] The final updated detection results will be packaged and sent to the fusion unit 30 to ensure that the input data for subsequent illegal behavior analysis or trajectory tracking is comprehensive and correct.

[0191] Through this collaborative correction process, the system achieves more consistent and accurate recognition of the same vehicle or frame at the intersection of adjacent partitions, avoiding conflicting judgments or duplicate recognitions caused by varying thresholds or occlusion levels within the partition. Furthermore, the introduction of confidence weighting and cross-validation mechanisms fully leverages the complementary information from different cameras within adjacent partitions, improving the reliability and accuracy of multi-camera fusion monitoring in complex environments, thereby providing more robust data support for subsequent traffic violation identification and early warning.

[0192] In a further embodiment of this application, multi-source traffic data includes not only radar detection, meteorological data, and vehicle-to-vehicle network information, but also road geometry information for the target road area. Road geometry information refers to precise descriptions or vectorized map data of road alignment, lane distribution, road curvature, number of lanes, slope, speed limit signs, and other information. This information can come from a geographic information system (GIS), high-precision mapping services, or a traffic management department's road network database.

[0193] By introducing road geometry information, the system can accurately align and compare the vehicle trajectory with the actual road layout, so as to better judge whether there are behaviors such as speeding, illegal lane changes, driving in the wrong direction, occupying the emergency lane, etc.

[0194] In practice, repeated video information can be used to establish a vehicle trajectory baseline. Within an area of ​​overlapping camera coverage, if a target vehicle is detected in different camera images with a similarity score greater than or equal to a first threshold, or if the vehicle is confirmed to be the same through collaborative correction of adjacent partitions, the relevant image frames are marked as repeated video information. This repeated video information provides a stable and highly consistent imaging basis and can be considered a mainline trajectory for cross-camera tracking.

[0195] The system can confirm the consistency of a target vehicle at the intersection of multiple cameras' coverage based on timestamps, vehicle appearance characteristics, or external Internet of Vehicles information, thereby forming a continuous record of the vehicle's travel path. This process effectively reduces the uncertainty caused by obstructed vision or other noise from a single camera, making repeated video information a reliable benchmark for vehicle trajectory analysis.

[0196] In some scenarios, the vehicle may be photographed from different perspectives, or be affected by factors such as partial occlusion and lighting changes, making it difficult for a single camera to continuously identify it.

[0197] However, if the similarity score is between the second threshold and the first threshold, and the difference analysis confirms that these image frames belong to complementary video information, it means that these frames can supplement the vehicle position from different perspectives or in a partially occluded state.

[0198] Based on this, the fusion unit 30 can superimpose the complementary video information to infer the target vehicle's complete or more complete temporal trajectory. For example, when a vehicle temporarily enters a ramp or blind spot and then returns to the main lane, the complementary video information can help reconstruct the missing path and reduce the possibility of vehicle loss.

[0199] After completing the fusion processing of repeated video information and complementary video information, a relatively complete vehicle trajectory data set is obtained, which includes information such as the vehicle's temporal position, speed, and the camera area it is located in.

[0200] Next, the system aligns this vehicle trajectory data with road geometry information. Specific steps may include:

[0201] If the vehicle trajectory is represented by pixel coordinates or camera coordinate system, it can first be projected or transformed into the global coordinate system or map coordinate system through the automatic calibration subunit (as described above);

[0202] Based on road geometry information, identify which lane or ramp the vehicle is in, and determine the relationship between its travel path and road lane lines, dividers, speed limit signs, etc.

[0203] By comparing the road's prescribed speed, curve radius, prohibited lane change section, emergency lane and other information, it is determined whether the vehicle trajectory is inconsistent with the prescribed traffic rules.

[0204] This comparison process generates a result that identifies whether the target vehicle has violated traffic regulations. For example, if a vehicle exceeds the speed limit threshold set by a speed limit sign on a certain road section, or crosses lane lines multiple times in a no-lane-change zone, the comparison result will indicate a potential violation.

[0205] Based on the comparison results, if the system determines that the target vehicle has indeed committed a preset violation (such as speeding, illegal lane changing, driving in the wrong direction, occupying the emergency lane, etc.), the system will automatically extract relevant vehicle identification information, such as license plate number, vehicle model, color, as well as the timestamp and geographic location of the violation.

[0206] At the same time, the system will generate corresponding violation indication information based on relevant regulatory or law enforcement needs. This indication information may include encrypted vehicle identity data, violation type, time and location, and screenshots or short video clips of the violation. This information is used to report to traffic law enforcement departments in real time or send to the traffic management backend for subsequent law enforcement or punishment.

[0207] Under this implementation, because the system introduces road geometry information and combines the advantages of repeated video information and complementary video information, it can more accurately locate the vehicle's status in the road space and maintain a complete grasp of the vehicle's trajectory in complex road conditions and multi-camera scenarios.

[0208] For management departments, this solution significantly improves the ability and accuracy of capturing various types of illegal behaviors, while reducing reliance on manual inspections and post-event video screenings, thereby achieving more efficient and intelligent road traffic supervision and early warning.

[0209] In this way, this application establishes a trajectory baseline based on image frames marked as repeated video information, combines image frames marked as complementary video information to improve the vehicle movement position, and then uses road geometry information to align and compare vehicle trajectory data to determine whether there is a specific traffic violation, and generates corresponding indication information after confirming the violation. This technical process not only improves the integrity of the multi-camera monitoring system in cross-camera tracking, but also significantly enhances the accuracy of violation detection and positioning.

[0210] See Figure 3 , Figure 3 This is a flowchart of a method for generating complementary video information provided in an embodiment of the present application; as an optional implementation, the generating of complementary video information includes S101 to S104, wherein:

[0211] S101: performing foreground extraction processing on the image frame to determine a suspicious target area;

[0212] S102: Matching the suspicious target area with the vehicle position and / or lane information in the multi-source traffic data to generate a difference parameter;

[0213] S103: In response to the difference parameter being greater than or equal to a preset determination parameter, marking the corresponding image frame as complementary video information;

[0214] S104: In response to the difference parameter being smaller than the preset determination parameter, updating the corresponding image frame to repeated video information.

[0215] In some cases, similarity scores alone cannot clearly determine whether an image frame is simply a repeated shot of the same scene, or a complementary shot where the perspective difference provides additional information. To this end, this application proposes a technical solution for performing foreground extraction processing on image frames. Specifically, when the similarity score is detected to fall between the second threshold and the first threshold (or other interval range specified by the system), further foreground / background separation is performed to identify suspicious target areas and perform difference analysis based on multi-source traffic data.

[0216] In a specific implementation, in the early image preprocessing stage, the detection unit 20 will use background subtraction, motion target detection or semantic segmentation methods based on deep learning for each frame of the image to extract the contour area of ​​the vehicle or other suspicious moving targets.

[0217] For example, commonly used background subtraction algorithms include Mixture of Gaussians (MOG2) and the KNN method. Deep learning methods can use models such as U-Net or Mask R-CNN to segment foreground and background. Based on this, the detection unit 20 can obtain several foreground regions, which are suspicious target areas and include important objects such as vehicles and pedestrians.

[0218] After extracting the suspicious target area, the detection unit 20 further compares it with externally sourced vehicle position and lane information. Vehicle position can come from real-time GPS coordinates from the connected vehicle system, roadside radar detection data, or vehicle coordinates calculated by integrating calibrations from various cameras. Lane information includes specific lane numbers, lane marking locations, and current speed limits.

[0219] If the spatial position, shape, or motion trajectory of the foreground target is significantly inconsistent with the vehicle positioning data or corresponding lane information at that moment, it means that this image frame may provide a different perspective or new available information, which helps to improve the trajectory of the target vehicle in scenes with occlusion or poor perspective.

[0220] On the contrary, if the detection results of the foreground target under different cameras are highly consistent with the multi-source traffic data, it may mean that the frame only repeatedly captures the same vehicle and the same lane position, and there is no substantial difference from the content captured by other cameras.

[0221] During this process, the system calculates a difference parameter to quantify the degree of difference between the suspicious target area and the existing vehicle or lane information.

[0222] For example, if the license plate position, size, trajectory inference speed, etc. of the detected foreground target deviate significantly from the external data and there is no other reasonable explanation (such as measurement error, time synchronization deviation), the difference value will be high; if the two are highly consistent or well matched, the difference value will be very low.

[0223] When the difference parameter is greater than or equal to the preset judgment parameter threshold, it means that the suspicious target area provides information that is different from conventional or known data in time and space. Therefore, the corresponding image frame can be marked as complementary video information, which means that the vehicle position, outline or other elements are supplementarily identified from different perspectives, which can help the system understand the driving condition of the target vehicle more comprehensively.

[0224] If the difference parameter is less than the preset judgment parameter, it means that the frame is actually basically consistent with the scene captured by other cameras and no valid additional information is introduced. At this time, the system can update it to repeated video information to avoid the repeated use of subsequent calculation or analysis resources for the actually identical or similar pictures.

[0225] For example, during the nighttime period, a camera captures a vehicle changing lanes at a relatively long distance. Due to the poor lighting, the image similarity algorithm is temporarily unable to determine whether the vehicle captured by the other camera is merely an extension of the same frame. At this point, if foreground extraction reveals that the vehicle is near the lane line and its movement trend is slightly different from the GPS trajectory reported by the Internet of Vehicles, the difference parameter will be evaluated as high, which means that the frame can make up for the information missing due to the nighttime lighting difference and is therefore marked as complementary video information. Conversely, in cases of good lighting or when the vehicles are clearly in the same position and posture, the system detects a small difference and updates it to duplicate information to maintain data consistency and simplicity.

[0226] This approach, based on foreground extraction and difference parameter determination, ensures the system doesn't rely too heavily on a single similarity score, but rather more flexibly and accurately identifies images from different cameras or viewpoints that complement vehicle detection. Combined with the aforementioned mechanisms of dynamic threshold adjustment, automatic sub-unit calibration, and collaborative correction of adjacent partitions, the system can fully exploit complementary information in complex road environments, occluded scenes, or with significant viewpoint differences, further improving the overall accuracy and robustness of cross-camera vehicle recognition and violation detection.

[0227] Those skilled in the art will understand that in the above-described methods of specific embodiments, the order in which the steps are presented does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible inherent logic. It should be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.

[0228] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

Claims

1. Traffic incident violation warning system based on AI and multi-source data fusion, characterized by: The system comprises: an acquisition unit, which acquires video data collected by multiple cameras deployed in a target road area and multi-source traffic data provided by at least one external data source; A detection unit, based on the video data, determines camera images with overlapping coverage areas and generates a detection result; wherein the detection result includes repeated video information and complementary video information; a fusion unit, which performs fusion analysis on the detection result and the multi-source traffic data to generate traffic violation indication information; an early warning unit, generating an early warning signal based on the traffic violation indication information; wherein the early warning signal is used to indicate a potential and / or ongoing traffic violation event; Generating the detection result includes: Performing time stamp alignment on the video data collected by the multiple cameras to generate multiple image frame sets; For the image frame set, extract image features using a preset model to generate a feature vector for evaluating picture similarity; Determining, based on the feature vector, image similarities between multiple image frames in the same image frame set; In response to the picture similarity score being greater than or equal to a first threshold, marking the corresponding image frame as duplicate video information; performing difference analysis on the image frames whose picture similarity scores are greater than or equal to the second threshold and less than the first threshold, and generating complementary video information; Integrate the image frames marked as repeated video information and complementary video information to generate the detection result, and output the detection result to the fusion unit; The detection unit performs a first dynamic adjustment on the first threshold and the second threshold based on the multi-source traffic data; The multi-source traffic data includes: traffic volume, weather conditions and vehicle speed; The traffic volume is obtained through radar detection data, the weather conditions are obtained through meteorological data, and the vehicle speed is obtained through Internet of Vehicles information; The detection unit is further configured as: Based on the detection result, the target road area is divided into a plurality of camera partitions; wherein the cameras corresponding to the image frames whose picture similarity scores are greater than or equal to a second threshold are considered to be in the same partition; In each of the camera zones, calculating a zone environment complexity index of the camera zone based on local multi-source traffic data; Based on the partition environment complexity index, performing a second dynamic adjustment on the first threshold and the second threshold; The first threshold and the second threshold of different camera partitions are independent of each other.

2. The traffic incident violation warning system based on AI and multi-source data fusion according to claim 1 is characterized in that: The detection unit also includes: The automatic calibration subunit obtains the spatial mapping relationship of the multiple cameras based on the feature points of the target road area, and performs spatial correction on the video data before calculating the picture similarity.

3. The traffic incident violation warning system based on AI and multi-source data fusion according to claim 2 is characterized in that: The detection unit is further configured as: When overlapping areas are detected between different partitions, the corresponding partitions are marked as adjacent partitions; For the adjacent partitions, the similarity determination and difference analysis results are collaboratively corrected.

4. The traffic incident violation warning system based on AI and multi-source data fusion according to claim 3 is characterized in that: The collaborative correction includes: Based on the degree of overlap of coverage areas between the adjacent partitions, performing confidence weighting processing on the similarity determination and the difference analysis results; In response to detecting that the same target vehicle and / or the same frame image generates a conflict determination in the adjacent partitions, recalculating the similarity score and / or the difference score through cross-validation; In response to the completion of the recalculation, the repeated video information and the complementary video information are marked as updated, and the updated detection result is output.

5. The traffic incident violation warning system based on AI and multi-source data fusion according to claim 4 is characterized in that: The multi-source traffic data further includes: road geometry information; the fusion analysis of the detection result and the multi-source traffic data to generate traffic violation indication information includes: Based on the image frames marked as repeated video information, the consistency of the target vehicle's continuous tracking trajectory in the overlapping area covered by multiple cameras is confirmed, and the repeated video information is used as the vehicle trajectory benchmark; Based on the image frames marked with complementary video information, the moving position of the target vehicle under different viewing angles or occlusion conditions is identified to generate vehicle trajectory data; Aligning and comparing the vehicle trajectory data with road geometry information to generate a comparison result; Based on the comparison result, determining whether the target vehicle has committed a preset illegal behavior; In response to the target vehicle having the preset illegal behavior, corresponding vehicle identification information and timestamp are extracted, and indication information for indicating the illegal behavior is generated.

6. The traffic incident violation warning system based on AI and multi-source data fusion according to claim 1 is characterized in that: Generating complementary video information includes: Performing foreground extraction processing on the image frame to determine a suspicious target area; Matching the suspicious target area with the vehicle position and / or lane information in the multi-source traffic data to generate a difference parameter; In response to the difference parameter being greater than or equal to a preset determination parameter, marking the corresponding image frame as complementary video information; In response to the difference parameter being smaller than the preset determination parameter, the corresponding image frame is updated to repeated video information.

Citation Information

Patent Citations

  • Automatic detection system for highway traffic event

    CN110570664A

  • Method for detecting traffic violation

    US20160034778A1