Monitoring event analysis method, system and equipment based on road video cloud networking
By dynamically generating event determination thresholds in roadside camera video streams and real-time monitoring of video quality, the performance degradation problem caused by the inability to continuously update the model in the prior art is solved, and high accuracy and robust traffic event recognition in complex traffic states are achieved.
Patent Information
- Application Number
- CN202510886871.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing computer vision-based video event recognition technology cannot be continuously updated after decoupling of model training and deployment, resulting in degradation of model performance, inability to adapt to changes in traffic state, prone to false alarms or missed alarms, and lack real-time monitoring of video quality, affecting the credibility of the identification results.
Video streams are collected through roadside cameras, combined with pre-trained visual models, dynamically generate traffic evidence ratio components, stay frame thresholds, corner mutation evidence ratio components, accident judgment time thresholds and restricted area residence time thresholds, and use historical correct detection rate and false alarm rate to make event judgments, and monitor video code rate and packet loss rate in real time to build an adaptive adjustment mechanism.
It realizes that in the long-term operation scenario after the model is cold-started, the accuracy and system robustness of traffic event recognition can be improved, and it can adapt to complex road conditions, identify multiple types of traffic abnormal events, and aggregate complete event records, which is convenient for unified scheduling and management of the system.
Smart Images

Figure CN120388336A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of intelligent transportation, video image processing and artificial intelligence analysis, and particularly relates to a method, system and device for analyzing highway video cloud networking monitoring events. Background Art
[0002] With the rapid development of urban intelligent transportation systems, road video surveillance data has become an important data source for traffic operation state analysis and event recognition. Especially in scenarios such as highways, urban expressways, and congestion-prone locations, using video streams for traffic event detection has become a key technical path to improve road safety and traffic efficiency.
[0003] Existing video event recognition technologies based on computer vision mainly rely on deep learning models to process video data collected by cameras to identify events such as traffic congestion, traffic accidents, and illegal vehicle intrusion. Such systems usually include links such as video acquisition, edge access, image recognition, trajectory tracking, behavior modeling, and event output. Among them, data including but not limited to, for example, video surveillance, log analysis, real-time bitrate detection, etc. are comprehensively statistically analyzed for various monitoring events, which can include the total number of events, the number of false alarms, the number of correct alarms, the proportion of various events, the ranking of the number of line events, etc. Event data from different subsystems can be collected by designing a unified data interface; big data analysis and machine learning algorithms are used to automatically identify and classify events, distinguish correct alarms from false alarms, and calculate various statistical indicators; a visual dashboard is constructed to display information such as statistical results, trend charts, and ranking tables in real time to support managers in making quick decisions.
[0004] However, in actual deployment, existing technologies have difficulty in decoupling model training and deployment. Existing technologies mostly rely on pre-trained visual models, which are fine-tuned using existing labeled data in the initial deployment phase to improve accuracy. However, after the system enters the large-scale operation stage, it is limited by computing resources and network transmission bandwidth, and the model parameters cannot be continuously updated, resulting in model performance degradation over time, especially when specific road sections, time periods or traffic conditions change, which is prone to false alarms or missed alarms. For example, the patent document with announcement number CN114738627B provides an intelligent traffic monitoring system and monitoring method, in which event judgment parameters are rigid and lack self-adaptation capabilities. For example, the vehicle speed threshold, angle deviation judgment standard, stay time threshold, etc. are mostly set as fixed empirical values, without considering the current recognition accuracy of the model, real-time traffic conditions, or changes in environmental factors, which can easily cause the rules to fail. For example, the same threshold performs very differently during the day and at night, and at high speeds and in congested conditions. However, general systems lack the ability to recognize the quality of their own input data. For example, the patent document with announcement number CN118247761B describes an on-board monitoring system for intelligent traffic inspection vehicles. Because traditional systems generally ignore the impact of transmission quality indicators such as video bit rate and network packet loss on recognition results, once the video quality degrades, even if the recognition result is abnormal, it is difficult for the system to identify whether it is a misjudgment due to inaccurate viewing, thereby affecting the credibility of the judgment result. Current technology is limited by the inability to continuously train models. The existing system lacks a mechanism that does not change the main model structure but can dynamically optimize based on historical judgment results and real-time recognition performance, often causing the system to fall into a dilemma where it cannot adjust parameters or retrain. Therefore, in order to achieve intelligent fusion and real-time statistics of cross-system data, the accuracy and visualization level of event monitoring should be comprehensively improved to provide decision support for intelligent monitoring and operation and maintenance management. Summary of the Invention
[0005] The purpose of the present invention is to propose a method, system and equipment for analyzing highway video cloud network monitoring events to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.
[0006] To achieve the above objectives, according to one aspect of the present invention, a method for analyzing highway video cloud network monitoring events is provided, the method comprising the following steps:
[0007] The video stream collected by the roadside camera is connected to the analysis system through the edge gateway. The video frame is accompanied by the camera number, geographic coordinates and timestamp;
[0008] The analysis system analyzes each frame of image, identifies the target object and calculates its moving speed and turning angle, and analyzes and outputs the historical correct detection rate and historical false alarm rate;
[0009] Based on the data of the video stream collected by the roadside camera, dynamically generate the traffic flow evidence ratio component, the stay frame threshold, the turning angle mutation evidence ratio component, the accident discrimination duration threshold, and the restricted area stay duration threshold: The traffic flow evidence ratio component is generated by combining the adjusted evidence ratio factor with the road speed limit and the historical free flow speed. Among them, the historical detection evidence ratio is calculated using the historical correct detection rate and the historical false alarm rate, and the adjusted evidence ratio factor is formed through the historical detection evidence ratio; the stay frame threshold is obtained by rounding up the product of the shortest continuous duration required for congestion determination and the video sampling rate; the turning angle mutation evidence ratio component is obtained by adding the mean value of the steering angle change to the combination of the turning angle mutation benchmark and the historical detection evidence ratio; the accident discrimination duration threshold is obtained by adding the congestion determination duration to the pre-set accident delay compensation duration; the restricted area stay duration threshold is obtained by dividing the physical length of the restricted area by the real-time average vehicle speed on the same road section.
[0010] Apply the object detection and tracking algorithm to each frame of the image, assign a unique identifier to each target object, and calculate the speed and angle changes in real time; for each target with an identifier, obtain its speed and the number of low-speed frames, and determine whether a traffic congestion event occurs by combining the traffic flow evidence ratio component and the stay frame threshold; obtain the single-frame steering angle change and speed change of each target with an identifier, and determine whether an accident event occurs by combining the turning angle mutation evidence ratio component and the traffic flow evidence ratio component; when the center of the bounding box of each target with an identifier enters the set restricted area and the continuous stay time reaches or exceeds the restricted area stay duration threshold, it is determined that an illegal intrusion event occurs.
[0011] Aggregate multiple recognition results with the same identifier, the same event type, and adjacent or overlapping in time to form a complete event record, and the event record includes the event type, target identifier, camera number, start and end times, and geographical coordinates.
[0012] Furthermore, the analysis system adopts a real-time detection network based on deep learning and a multi-target association tracking algorithm to realize the positioning and recognition of each frame of the target and the real-time calculation and annotation of its speed and angle change amount.
[0013] Furthermore, the stay frame threshold is dynamically calculated and rounded up based on the shortest congestion determination time and the actual sampling frame rate of the video stream; among them, the average flow rate is calculated based on the speeds of each vehicle in the data of the video stream collected by the roadside camera, and the vehicles with speeds lower than 10% of the average flow rate in the data of the video stream are judged as congested vehicles. The average value of the stay time of the congested vehicles is used as the shortest congestion determination time, and the frames with speeds lower than 10% of the average flow rate in the data of the video stream are used as low-speed frames, and the number of low-speed frames is used as the number of low-speed frames.
[0014] Further, the accident discrimination duration threshold is obtained by adding the shortest congestion determination time and the accident delay compensation time, where the accident delay compensation time is the average value of the time for each vehicle to stop moving and then resume moving in the video stream collected by the roadside camera.
[0015] Further, the prohibited area stay duration threshold is calculated based on the ratio of the physical length of the prohibited area to the real-time average driving speed of this section of the road, where the physical length of the prohibited area is the length obtained by proportionally scaling the set prohibited area in the physical space in the video stream or by actual measurement.
[0016] Further, the traffic flow certificate ratio component is generated by combining the adjusted certificate ratio factor with the road speed limit and the historical free flow speed. Specifically, it further includes:
[0017] Obtain the historical correct detection rate, historical false alarm rate, target detection rate, and target false alarm rate; where the historical correct detection rate is the frequency of detecting real congestion events within a period of time; the historical false alarm rate is the frequency of misjudging normal traffic as congestion within the same period of time.
[0018] The ratio of the historical correct detection rate to the historical false alarm rate is the historical detection certificate ratio, the natural logarithm function of the historical detection certificate ratio is the historical empirical certificate ratio, and the ratio of the historical empirical certificate ratio to one hundred is the adjusted certificate ratio factor.
[0019] Divide the historical false alarm rate by the historical detection certificate ratio to obtain the target false alarm rate, and the part obtained by subtracting the target false alarm rate from the whole one is the target detection rate.
[0020] Select a sampling time window, obtain the correct detection rate and false alarm rate within the sampling time window. If the value of the correct detection rate is less than the target detection rate, then update the specific value of the adjusted certificate ratio factor by multiplying the specific value of the adjusted certificate ratio factor by the ratio of the historical correct detection rate divided by the target detection rate.
[0021] Obtain the road speed limit of the section to be detected, and obtain the arithmetic average of the vehicle flow speeds of the section to be detected in the past period of time, which can be multiple different past sampling time windows, as the historical free flow speed. Combine the minimum value of the obtained road speed limit and the historical free flow speed with the updated adjusted certificate ratio factor to generate the driving speed threshold.
[0022] Further, the corner mutation certificate ratio component is obtained by adding the mean value of the statistical steering angle change to the combination of the historical detection certificate ratio and the corner mutation reference. Specifically, it further includes:
[0023] The maximum value in the sequence of corner difference within the sampling time window is the peak value of corner difference;
[0024] Calculate the ratio of the standard deviation of the corner sequence to the mean value of the corner sequence as the sequence corner coordinate, multiply the sequence corner coordinate by the peak value of corner difference to obtain the corner mutation reference, and add the mean value of the corner sequence to the combination of the historical detection ratio and the corner mutation reference to obtain the corner mutation certification component;
[0025] Among them, the historical detection ratio and the corner mutation reference are combined in a multiplicative manner.
[0026] Furthermore, it also includes that during the stage when the video stream collected by the roadside camera accesses the analysis system through the edge gateway, the average bit rate and network packet loss rate of the video stream are monitored in real time to obtain the real-time packet loss rate, and the video quality is marked as abnormal and uploaded as a special event through the lower limit of the acceptable bit rate ratio and the upper limit of the packet loss rate ratio; among them, the positive reporting rate of multiple different sampling time windows is continuously monitored and the mean value of the positive reporting rate is obtained. The lower limit of the acceptable bit rate ratio is automatically taken as the product of the average bit rate and the mean value of the positive reporting rate according to historical statistics, the upper limit of the packet loss rate ratio is the product of the average packet loss rate and the historical detection ratio, and the average packet loss rate is the arithmetic mean of the real-time packet loss rate.
[0027] The present invention also provides a highway video cloud networking monitoring event analysis system, which includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it realizes the steps in the highway video cloud networking monitoring event analysis method. The highway video cloud networking monitoring event analysis system can run on computing devices such as desktop computers, laptop computers, palm computers, and cloud data centers. The operable system may include, but is not limited to, a processor, a memory, and a server cluster. The processor executes the computer program and runs in the following system units:
[0028] A data acquisition unit, which is used for the video stream collected by the roadside camera to access the analysis system through the edge gateway, and attaches data such as camera number, geographical coordinates, and timestamps to each frame of the video stream. The analysis system conducts analysis, identifies the moving speed and angle of the vehicle in each frame, and analyzes and outputs the historical correct detection rate and historical false alarm rate;
[0029] A variable operation unit, which is used to dynamically generate traffic certification components, stay frame thresholds, corner mutation certification components, accident discrimination duration thresholds, and restricted area stay duration thresholds based on the data of the video stream collected by the roadside camera;
[0030] A target recognition unit for applying a real-time target detection and tracking algorithm to each frame of image, assigning a unique identifier to each pedestrian or vehicle, and calculating its instantaneous speed and the change in steering angle between adjacent frames;
[0031] A target judgment unit for, for each target with an identifier, when its speed continuously drops below the traffic flow proof component and the number of low-speed frames reaches or exceeds the stay frame threshold, determining that a traffic congestion event has occurred; when the change in the steering angle of a single frame exceeds the corner mutation proof component and its speed remains below the traffic flow proof component within the accident discrimination duration threshold, determining that an accident event has occurred; for the center of the bounding box of each target with an identifier, when the center of its bounding box enters a set prohibited area and the continuous stay time reaches or exceeds the prohibited area stay duration threshold, determining that an illegal intrusion event has occurred;
[0032] A judgment transmission unit for aggregating multiple discrimination results with the same identifier, the same event type, and adjacent or overlapping in time among the recognized results, and outputting them as a complete event record, which includes the event type, target identifier, camera number to which it belongs, start and end times, and geographical coordinates.
[0033] Correspondingly, the present invention also provides an electronic device, a readable storage medium, and a computer program product:
[0034] An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0035] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0036] A computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0037] The beneficial effects of the present invention are as follows: The present invention relates to an event analysis method, system and device based on highway video cloud networking monitoring. It can collect video streams through roadside cameras and, in combination with a pre-trained vision model, perform real-time identification and tracking of vehicle speed, steering angle, etc. The system dynamically generates traffic flow evidence score components, stay frame thresholds, corner mutation evidence score components, accident discrimination duration thresholds and restricted area stay duration thresholds according to historical correct detection rates and false alarm rates, which are used for the determination of congestion, accident and illegal intrusion events respectively. It also introduces a video bit rate and packet loss rate monitoring mechanism to mark and report abnormal video quality. This method has an adaptive adjustment ability and is applicable to long-term operation scenarios after model cold start, which can effectively improve the accuracy of traffic event identification and the robustness of the system.
[0038] The present invention realizes an external control mechanism for behavior accuracy after model cold start. By introducing an adjustment evidence ratio factor and combining road speed limits and historical free flow speeds, it adaptively generates traffic flow evidence score components for dynamic discrimination of events such as traffic congestion and accidents. This design breaks through the performance degradation problem caused by the inability to update model parameters after the cold start stage, and realizes the self-learning and self-optimization of external judgment strategies on the premise that the main model is fixed. The present invention also constructs a corner mutation evidence score component, and combines the average value, standard deviation and peak data of angle changes in the real-time sampling window to achieve accurate identification of abnormal vehicle offset behaviors, effectively improving the response ability to sudden events such as traffic accidents. It can effectively distinguish normal lane-changing behaviors from unexpected steering behaviors and enhance the discrimination robustness of the system in actual complex road conditions. In the highway video cloud networking environment, it can simultaneously identify multiple types of traffic abnormal events such as congestion, accident and illegal intrusion, and has the ability to aggregate repeated judgment results into complete event records, which is convenient for system unified scheduling and record management, and is applicable to large-scale deployment and deep integration in intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] By elaborating on the embodiments shown in the accompanying drawings in detail, the above and other features of the present invention will become more obvious. The same reference numerals in the drawings of the present invention represent the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0040] Figure 1 It shows a flowchart of an event analysis method based on highway video cloud networking monitoring;
[0041] Figure 2 It shows a system structure diagram of an event analysis system based on highway video cloud networking monitoring. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The concept, specific structure, and technical effects of the present invention will be clearly and completely described below in combination with embodiments and the accompanying drawings to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0043] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is two or more, "greater than", "less than", "exceeding", etc. are understood not to include the present number, and "above", "below", "within", etc. are understood to include the present number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0044] As Figure 1 shown is a flowchart of a method for analyzing highway video cloud networking monitoring events according to the present invention. The following will be combined with Figure 1 to elaborate on a method, system, and device for analyzing highway video cloud networking monitoring events according to an embodiment of the present invention.
[0045] The present invention proposes a method for analyzing highway video cloud networking monitoring events. The method specifically includes the following steps:
[0046] The video stream collected by the roadside camera is accessed to the analysis system through the edge gateway, and data such as the camera number, geographical coordinates, and timestamp are attached to each frame in the video stream. The analysis system performs analysis, analyzes and identifies the moving speed and angle of the vehicle in each frame, and analyzes and outputs the historical correct detection rate and historical false alarm rate.
[0047] Based on the data of the video stream collected by the roadside camera, the traffic flow evidence score component, the stay frame threshold, the turning angle mutation evidence score component, the accident discrimination duration threshold, and the restricted area stay duration threshold are dynamically generated.
[0048] Apply real-time object detection and tracking algorithms to each frame of the image, assign a unique identifier to each pedestrian or vehicle, and calculate its instantaneous speed and the change in steering angle between adjacent frames.
[0049] For each tagged target, when its speed continuously drops below the traffic flow evidence score component and the number of low-speed frames reaches or exceeds the stay frame threshold, it is determined that a traffic congestion event has occurred; when the single-frame steering angle change exceeds the turning angle mutation evidence score component and its speed remains below the traffic flow evidence score component within the accident discrimination duration threshold, it is determined that an accident event has occurred; for the center of the bounding box of each tagged target, when the center of its bounding box enters the set restricted area and the continuous stay time reaches or exceeds the restricted area stay duration threshold, it is determined that an illegal intrusion event has occurred.
[0050] Aggregate multiple discrimination results for the same identified label, the same event type, and adjacent or overlapping in time, and output them as a complete event record, which includes the event type, the target label, the camera number to which it belongs, the start and end times, and the geographical coordinates.
[0051] Among them, the traffic flow certificate ratio component is generated by combining the adjusted certificate ratio factor with the road speed limit and the historical free flow speed. The stay frame threshold is obtained by rounding up the product of the shortest duration required for congestion determination and the video sampling rate. The turning angle mutation certificate ratio component is obtained by adding the mean value of the statistical turning angle change to the combination of the turning angle mutation reference and the historical detected certificate ratio. The accident discrimination duration threshold is obtained by adding the congestion determination duration and the pre-set accident delay compensation duration. The restricted area stay duration threshold is obtained by dividing the physical length of the restricted area by the real-time average vehicle speed of the same road section.
[0052] In some embodiments, the streaming video access can be through the H264 video stream collected by the roadside camera and accessed to the system through the edge gateway, and the gateway pushes video frames to the cloud or regional nodes. Each frame carries metadata such as camera number, GPS coordinates, timestamp, etc. Each frame of the image is processed by real-time detection models including but not limited to YOLO-V5 and / or other real-time detection models to identify all pedestrians and vehicles, and a unique ID is assigned to each target through tracking algorithms such as DeepSORT. At the same time, the instantaneous speed of the target based on the position difference between adjacent frames and the steering angle difference of the target between this frame and the previous frame are calculated. Among them, a pre-trained large vision model is embedded in the analysis system, including but not limited to large vision models such as YOLO. The pre-trained large vision model can be used to identify and classify event behaviors according to the data of the video stream collected by the roadside camera. In the initial stage of the operation of the analysis system, there should be a cold start phase. In the cold start phase, the pre-trained large vision model identifies and classifies event behaviors according to the existing training data of the video stream, and counts the correct rate obtained from the identification and classification of event behaviors as the historical correct detection rate, and the error rate obtained as the historical false alarm rate, and performs fine-tuning of the model, that is, the training update of some neural networks. The existing training data of the video stream can be training data in the form of a training set and a validation set, which are marked with the answers for verifying the identification and classification of event behaviors as standards. For example, the values of recall and precision are selected in the process of the model calculating the F1-score. Thus, it is convenient for the pre-trained large vision model to obtain the historical correct detection rate and historical false alarm rate according to the data of the video stream collected by the roadside camera in the cold start phase. After the cold start phase is the normal operation phase. According to the normal operation phase of the existing technology, if the pre-trained large vision model still needs to be fine-tuned, the computing power cost will be very high, and it is difficult to support the data of so many video streams on the road to always update the model parameters, and the data cluster will bear great pressure. However, if the parameters are not updated, the recognition efficiency and accuracy are not enough, so it often makes technicians in this field in a dilemma.
[0053] After entering the normal operation phase, due to the huge amount of highway video data and limited computing power, it is impossible to continuously update the model parameters, resulting in a possible decline in the recognition accuracy. The system needs to rely on the statistical values in the cold start period as the subsequent judgment criteria. However, once these static data deviate from the actual situation, the model robustness will decline sharply.
[0054] Calculate the displacement of the vehicle in the video stream along with the video playback time and its displacement, and scale it proportionally to the actual physical distance according to the lens shooting ratio to obtain the actual speed and actual moving distance of the vehicle in the video stream, so as to obtain the instantaneous speed of the vehicle in each frame of the video stream.
[0055] Among them, the identifier is the annotation information of vehicle targets in each frame of the data of the video stream, and the event type is the text information describing the events detected for the vehicle targets in each frame of the data of the video stream. The data of the event type, target identifier, the camera number to which it belongs, start and end times, and geographical coordinates are data types that can be stored in the database. The respectively generated "congestion", "accident", and "intrusion" events should be understood as the event recognition classifications of the road video stream, and any other equivalent replacements fall within the protection scope of the embodiments described in the present invention.
[0056] Further, the analysis system adopts a real-time detection network based on deep learning and a multi-target association tracking algorithm to achieve the real-time calculation and annotation of the positioning and recognition of each frame target and its speed and angle change amounts.
[0057] In some embodiments, the number of frames in the data of the video stream collected by the roadside camera is n, and the sequence number of the frames arranged in chronological order is t. For two frames, target recognition and target tracking can be performed through algorithm models of machine vision, etc. Each frame target is in each frame. The coordinates of the frame with the detection position and sequence number t are recorded as (x1 = 250px, y1 = 400px), and frame t + 1 = (x2 = 255px, y2 = 404px). Among them, the pixel scaling factor is 0.1 m / px, and the real-time inter-frame interval is 1 / 5fps = 0.2s. Obtain and record the instantaneous bit rate and network packet loss rate of each video stream.
[0058] Then, dynamically calculate that the displacement between the two coordinates is approximately equal to 0.64 m, and its instantaneous speed can be obtained by dividing 0.64 m by 0.2 s to get 3.2 m / s. Among them, record the change value of the angle of the moving direction between the coordinates in the two frames. The main axis of the lane direction can be corrected according to the camera angle, and the main axis direction can be obtained in real time. The difference between the main axis directions of the two frames is the turning angle, and the value of the turning angle is the single-frame steering angle change for subsequent dynamic discrimination.
[0059] The analysis system can also adopt a real-time detection network YOLOv5 based on deep learning and multi-object association tracking algorithms such as MHT and / or PDA to perform image recognition on the images of each vehicle in each frame. It can include but is not limited to using algorithms such as the ground-truth bounding box algorithm to identify the bounding boxes of each vehicle image in the frame, and the geometric center of the bounding box of each vehicle image is the center of its bounding box. For example, at a certain moment frame t, the center of the target bounding box moves from (x = 250px, y = 400px) to (x = 255px, y = 402px) in frame t+1. Given a camera field-of-view conversion factor of 0.1 m / px, the distance change is approximately equal to 0.54m. With an inter-frame time of 1 / 6s, the instantaneous speed is approximately 0.54 / (1 / 6) = 3.24m / s, and the steering angle is calculated to be 12° through the change in the main axis direction of the bounding box.
[0060] In some embodiments, for the calculation method of the positive detection rate and false alarm rate, data can be retrieved from an external database and / or by identifying and analyzing the video stream collected by roadside cameras, etc. Event records are detected from a time sampling window, including 7 records of "congestion", 3 records of "accident", and 2 records of "intrusion" identified. After verification, 8 positive detections and 4 false alarms are detected. Dynamically calculate the positive detection rate as 8 divided by (7 + 3 + 2) which is approximately equal to 66.7%, and the false alarm rate is approximately 33.3%.
[0061] Further, the stay frame threshold is dynamically calculated based on the shortest congestion determination time and the actual sampling frame rate of the video stream and rounded up.
[0062] Among them, the average flow rate is calculated based on the speeds of each vehicle in the data of the video stream collected by the roadside camera. Vehicles with speeds lower than 10% of the average flow rate in the video stream data are judged as congested vehicles. The average value of the stay time of the congested vehicles is used as the shortest congestion determination time, and the frames with speeds lower than 10% of the average flow rate in the video stream data are used as low-speed frames, and the number of low-speed frames is used as the number of low-speed frames.
[0063] In some embodiments, the shortest congestion determination time can be automatically calculated through traffic monitoring to be 2.5 s. The specific size of the actual sampling frame rate of the video stream can drop from the default frame rate of 6 fps to, for example, 5 fps when there is network congestion to ensure consistent time judgment accuracy under different frame rate conditions. The stay frame threshold is obtained by dynamically multiplying the shortest congestion determination time and the actual sampling frame rate of the video stream. When 2.5s is multiplied by 5fps, the stay frame threshold is 13 frames, and if the frame rate is restored to 6 fps, the stay frame threshold is 15 frames.
[0064] Further, the accident discrimination duration threshold is obtained by adding the shortest congestion determination time and the accident delay compensation time, where the accident delay compensation time is the average value of the time for each vehicle to stop moving and then resume moving in the video stream collected by the roadside camera.
[0065] In some embodiments, the accident delay compensation duration can be automatically obtained according to the data in the video stream collected by the roadside camera with an average stagnation of 4 s.
[0066] Further, the prohibited area stay duration threshold is calculated based on the ratio of the physical length of the prohibited area to the real-time average driving speed of this section, where the physical length of the prohibited area is the length obtained by equally scaling or actually measuring the set prohibited area in the physical space in the video stream.
[0067] In some embodiments, the physical length of the prohibited area can be statistically obtained through map measurement and is preferably set to 15 m.
[0068] The real-time average driving speed of this section is the real-time average vehicle speed of the same section in the video. The real-time average vehicle speed can be obtained according to real-time traffic flow statistics and is preferably 12 m / s. The prohibited area stay duration threshold is the physical length of the prohibited area, 15 m, divided by the real-time average driving speed of this section, 12 m / s, which is approximately equal to 1.25 s; if it is identified in the video stream that the sudden increase in traffic flow causes the average speed to drop to 10 m / s, then the prohibited area stay duration threshold is approximately 1.5 s. To adaptively determine the illegal stay determination conditions under different speed limits and different sections.
[0069] Further, among them, the traffic flow certificate ratio component is generated by combining the adjusted certificate ratio factor with the road speed limit and the historical free flow speed, and specifically further includes:
[0070] Obtain the historical correct detection rate, historical false alarm rate, target detection rate, and target false alarm rate, where
[0071] The historical correct detection rate is the frequency of detecting real congestion events in the past period of time;
[0072] The historical false alarm rate is the frequency of misjudging normal traffic as congestion in the same past period of time;
[0073] The ratio of the historical correct detection rate to the historical false alarm rate is the historical detection certificate ratio, the natural logarithm function of the historical detection certificate ratio is the historical empirical certificate ratio, and the ratio of the historical empirical certificate ratio to one hundred is the adjusted certificate ratio factor;
[0074] Divide the historical false alarm rate by the historical detection certificate ratio to obtain the target false alarm rate, and the part obtained by subtracting the target false alarm rate from the whole one is the target detection rate.
[0075] Select a sampling time window, obtain the correct detection rate and false alarm rate within the sampling time window. If the value of the correct detection rate is less than the target detection rate, then adjust the specific value of the adjustment proof ratio factor by multiplying it by the ratio obtained by dividing the historical correct detection rate by the target detection rate to update the specific value of the adjustment proof ratio factor;
[0076] Obtain the road speed limit of the road section to be detected, and obtain the arithmetic mean of the vehicle flow speeds in the past period (which can be multiple different sampling time windows in the past) of the road section to be detected as the historical free flow speed. Combine the minimum value of the road speed limit and the historical free flow speed with the updated adjustment proof ratio factor to generate a driving speed threshold.
[0077] In Embodiment 2, the road speed limit of 90 km / h is converted to 25 m / s, the historical free flow speed is 18 m / s, and the safety factor is 2%. Multiply the smaller of the two, 18 m / s, by 2% to obtain a driving speed threshold of 0.36 m / s. However, there are still some problems with such an implementation method, so a preferred embodiment is proposed.
[0078] In the preferred embodiment, the safety factor determines the mapping ratio of the traffic flow proof ratio component to the road speed limit or the free flow speed. If it is too high, congestion may be missed during detection; if it is too low, false alarms may occur easily.
[0079] Therefore, obtain the historical correct detection rate, historical false alarm rate, target detection rate, and target false alarm rate. Among them, the historical correct detection rate is the frequency of detecting real congestion events in the past period;
[0080] The historical false alarm rate is the frequency of misjudging normal traffic as congestion in the same past period;
[0081] Among them, the past period can be a set of multiple different sampling time windows, and the duration of one sampling time window is approximately 5 s to 10 s.
[0082] For example, the recently statistically obtained real congestion detection rate is 85% of the historical correct detection rate, and the recently statistically obtained real congestion detection false alarm rate is 15% of the historical false alarm rate.
[0083] The target detection rate and the target false alarm rate correspond to the historical correct detection rate and the historical false alarm rate. The target detection rate is higher than the historical correct detection rate, and the target false alarm rate is lower than the historical false alarm rate.
[0084] Specifically, the ratio of the historical correct detection rate to the historical false alarm rate is the historical detection certificate ratio. The natural logarithm function of the historical detection certificate ratio is the historical verification ratio. The ratio of the historical verification ratio to 100 is the adjusted certificate ratio factor. For example, calculating Ln(the historical detection certificate ratio) gives approximately 1.7, then the historical verification ratio to 100 gives 1.7%.
[0085] Then, divide the historical false alarm rate by the historical detection certificate ratio to obtain the target false alarm rate. The part obtained by subtracting the target false alarm rate from the integer 1 is the target detection rate. For example, the historical detection certificate ratio is 1.7, the historical false alarm rate is 15%, dividing the historical false alarm rate by the historical detection certificate ratio gives the target false alarm rate which can be rounded to 9%, then the target detection rate is 1 - 9% = 91%.
[0086] Furthermore, after the first calculation of the traffic flow certificate component, it is dynamically adjusted according to the historical road speed limit and the average vehicle speed during non-peak hours of the corresponding sections of each camera. Select the most recent sampling time window, which may not belong to the past period of time. Obtain the true correct detection rate and the true false alarm rate within the most recent sampling time window. If the value of the true correct detection rate is less than the target detection rate, then update the specific value of the adjusted certificate ratio factor by multiplying the specific value of the adjusted certificate ratio factor by the ratio of the historical correct detection rate divided by the target detection rate. For example, when the true false alarm rate within the most recent sampling time window has increased compared to its previous sampling time window, update the specific value of the adjusted certificate ratio factor of 1.7% by multiplying it by the ratio of 0.93 obtained by dividing the historical correct detection rate of 85% by the target detection rate of 9%. The specific value of the adjusted certificate ratio factor is updated to 1.6%. Among them, preferably, it is updated at any time by multiplying the smaller of the certificate ratio factor and the two. Obtain the road speed limit of the section to be detected, and obtain the arithmetic mean of the vehicle flow speeds of the past period of time for the section to be detected, which can be multiple different past sampling time windows, as the historical free flow speed. Combine the minimum value of the road speed limit and the historical free flow speed with the updated adjusted certificate ratio factor to generate the driving speed threshold.
[0087] The present invention calculates the historical detection certificate ratio by introducing statistical indicators such as the historical correct detection rate and the historical false alarm rate, and derives the historical verification ratio through the natural logarithm function of this ratio. Then, it combines with the percentage constant to form the adjusted certificate ratio factor. Finally, the adjusted certificate ratio factor is combined with the smaller value of the road speed limit or the historical free flow speed to generate the traffic flow certificate component for event recognition, and this ratio is dynamically adjusted according to the positive / false alarm rate within the actual sampling window.
[0088] Traditional traffic event judgment models based on speed thresholds often rely on manual experience to set static parameters. They are unable to adapt to changes in model accuracy under different road sections, different time periods, and different traffic conditions. They are unable to automatically adjust to fluctuations in recognition errors, are prone to insufficient generalization or overfitting problems, and lack a compensation mechanism for the degradation of model accuracy during continuous operation.
[0089] The technical solution described in the present invention introduces an adjustment ratio primer as a quantitative regulator for model accuracy control, and combines historical discrimination performance to construct an adjustable speed threshold through the calculation of the positive alarm rate / false alarm rate. This design takes into account both the modeling capability of the dynamic volatility of the data set itself and the closed-loop feedback control of the recognition accuracy index. The concept of information entropy can be introduced through the natural logarithm function to effectively smooth the fluctuation of accuracy response. This mechanism is equivalent to giving the event discrimination threshold a self-learning and self-adaptation capability, which enables the system to self-adjust the performance of different accuracy stages. It significantly improves the judgment accuracy under different traffic conditions, congestion, high speed, and abnormalities, and reduces the proportion of false alarms and missed alarms of the system in long-term operation. It reduces the frequency of manual intervention in parameter setting, improves the level of intelligent autonomy of the system, provides a quantifiable and traceable error correction path, and enables the system to have dynamic strategy adjustment capabilities.
[0090] The method described in this invention can alleviate the problem of model sluggishness after parameter freezing. Using the historical detection-to-evidence ratio obtained from a cold start as input, combined with real-time feedback to update and adjust the ratio, the decision threshold can be dynamically adjusted without retraining the main model. This creates a lightweight behavioral parameter relearning mechanism. Using the positive and false alarm rate data from the sampling window, it is equivalent to modifying the model's behavioral decisions without changing the model, resulting in pseudo-fine-tuning of the outer decision logic. This can also improve the model's long-term stability and adaptability, enable data-driven sustainable parameter migration, avoid performance solidification after a cold start, and resolve the engineering bottleneck of a dilemma.
[0091] Furthermore, the corner mutation verification ratio component is obtained by adding the mean of the statistical steering angle change to the historical detection verification ratio and the corner mutation benchmark, and specifically includes:
[0092] The maximum value in the rotation angle difference sequence within the sampling time window is the rotation angle difference peak value;
[0093] Calculate the ratio of the turn sequence standard deviation to the turn sequence mean as the sequence turn coordinate, multiply the sequence turn coordinate by the turn difference peak to obtain the turn mutation benchmark, and combine the turn sequence mean plus the historical detection certificate ratio with the turn mutation benchmark to obtain the turn mutation certificate ratio component;
[0094] The historical detection certificate ratio is combined with the corner mutation benchmark in a multiplicative manner.
[0095] In the fourth embodiment, with a sampling time window of 5 seconds, the sequence of angular difference within this 5-second sampling time window can be obtained by using the OpenCV module and the ultralytics module to identify the difference in the angle of vehicle movement in each frame relative to its angle in the previous frame. This sequence includes 8 frames, namely: 4°, 6°, 3°, 7°, 5°, 8°, 2°, 6°;
[0096] Calculate that the mean of the angular sequence of the angular difference sequence within this 5-second sampling time window is approximately equal to 5.1°, and its standard deviation of the angular sequence is approximately equal to 2.1°;
[0097] In some embodiments, the angular mutation proof ratio component is obtained by adding twice the standard deviation of the angular sequence to the mean of the angular sequence. However, there are still some problems with such an implementation method, and thus a preferred embodiment is proposed.
[0098] In the preferred embodiment,
[0099] Obtain the maximum value 8° in the angular difference sequence within this 5-second sampling time window as the angular difference peak value;
[0100] Calculate the ratio 0.41 of the standard deviation 2.1° of the angular sequence to the mean 5.1° of the angular sequence as the sequence angular coordinate. Multiply the sequence angular coordinate 0.41 by the angular difference peak value 8° to obtain the angular mutation reference 3.28°. The sum of the mean 5.1° of the angular sequence and the combination of the historical detection proof ratio and the angular mutation reference gives the angular mutation proof ratio component 10.7. Among them, the historical detection proof ratio 1.7 and the angular mutation reference 3.28° are combined in a multiplicative manner to be 5.6°.
[0101] The present invention calculates the sequence angular coordinate through the mean and standard deviation of the angular change within the sampling time window, and then multiplies it by the angle peak value to construct the angular mutation reference. Then, this reference is combined with the historical detection proof ratio to finally generate the angular mutation proof ratio component, which is used for the identification of events such as abnormal deflection (e.g., accidents).
[0102] In traditional traffic event recognition systems based on target trajectory analysis, natural lane changes or obstacle avoidance by vehicles often lead to misjudgments, especially when the camera angle is poor or there is data jitter. Generally, the system is unable to effectively identify whether the deviation is abnormal or whether the deflection is sudden, lacking the ability of statistical modeling for angle fluctuations and unable to quantitatively judge critical deflections.
[0103] The technical solution of the present invention constructs a dynamic corner offset index system. By using the ratio of the peak value to the standard deviation, an abnormal deviation judgment benchmark is established. Combining this benchmark with the ratio of the historical determination ability index forms an angle mutation threshold after risk adjustment. Through the above logic, the accurate identification of abnormal angle changes is realized, enhancing the robustness of the system in dealing with irregular trajectory behaviors. It can significantly reduce false judgments in normal offset behaviors such as in curves and lane changes. It can also effectively identify non-linear deflection behaviors caused by accidents or sharp turns, improving the recognition sensitivity of the system to the incoherence of the trajectory or the precursors of abnormal behaviors.
[0104] In the case where the system perspective is fixed after cold start and cannot adapt to the angle characteristics of the new road section, the technical solution of the present invention provides a self-learning angle behavior calibration mechanism. Based on the dynamic collection of corner sequences for each road section, an angle fluctuation feature space is formed. Without updating the model structure, it can adapt to terrain and streamline changes. Using local statistics such as angle mean, standard deviation, and peak value to construct a dynamic safety belt, strengthening the response ability to mutation behaviors such as accidents, and making up for the learning blind area of non-standard behaviors caused by insufficient coverage of the cold start data set. This can endow the system with a behavioral cognitive boundary, similar to using mathematical methods to delimit the behavioral boundary of whether it is a normal deflection or an accident for the existing model, realizing policy-level adjustment without retraining.
[0105] Further, preferably, it also includes that at the stage when the video stream collected by the roadside camera accesses the analysis system through the edge gateway, the average bit rate and network packet loss rate of the video stream are monitored in real time to obtain the real-time packet loss rate. Through the lower limit of the acceptable bit rate ratio and the upper limit of the packet loss rate ratio, the abnormal video quality is marked and then uploaded as a special event; where the positive reporting rate of multiple different sampling time windows is continuously monitored and the mean value of the positive reporting rate is obtained. The lower limit of the acceptable bit rate ratio is automatically taken as the product of the average bit rate and the mean value of the positive reporting rate according to historical statistics. The upper limit of the packet loss rate ratio is the product of the average packet loss rate and the historical detection ratio. The average packet loss rate is the arithmetic mean of the real-time packet loss rate.
[0106] In some embodiments, with a 10-second sampling time window, the sampling bit rates within this 10 s sampling time window are: 1600, 1400, 1500, 1300, 1700 units of kbps; the real-time packet loss rates are: 2%, 6%, 4%. During the process of the video stream accessing the analysis system through the edge gateway, that is, during the video access stage, quality detection is performed to obtain and record the instantaneous bit rate and network packet loss rate of each video stream. Preferably, the system's ping command can be called through the subprocess library, and according to the output result of the obtained ping command, the output result is parsed to calculate the packet loss rate.
[0107] Continuously monitor the positive reporting rate of multiple different sampling time windows and obtain the mean value of the positive reporting rate. The lower limit of the acceptable code rate ratio can be automatically taken as the product of the average code rate and the mean value of the positive reporting rate based on historical statistics;
[0108] The upper limit of the packet loss rate ratio can be set as the product of the average packet loss rate and the historical detection ratio, and the average packet loss rate is the arithmetic mean of the real-time packet loss rate.
[0109] In one embodiment, the calculated mean value of the positive reporting rate is 75%. The product of the average code rate of 1500 kbps and the mean value of the positive reporting rate of 75% is 1687.5 kbps; the combination of the arithmetic mean of the real-time packet loss rate of 4% and the historical detection ratio of 1.7 is 6.8%.
[0110] When the sampling code rate within the sampling time window is lower than the lower limit of the acceptable code rate ratio, then trigger to mark it as a "code rate insufficient" event; if the packet loss rate rises above the upper limit of the packet loss rate ratio, then trigger a "poor connection" event; upload and save the video data within the sampling time window that contains the "code rate insufficient" event and the "poor connection" event.
[0111] In the stage of the roadside camera video access analysis system of the present invention, the detection of the real-time code rate and packet loss rate of the video is introduced, and the lower limit of the code rate and the upper limit of the packet loss are constructed by combining the average code rate with the mean value of the positive reporting rate and the average packet loss rate with the historical ratio. When the video quality parameters exceed the above limit range, it is judged as an abnormal event such as insufficient code rate or poor connection and reported.
[0112] Existing event recognition systems generally ignore that low code rate or high packet loss leads to blurred targets and recognition failures, but the system cannot distinguish whether it is due to recognition incompetence or signal loss. Moreover, the network quality fluctuation is uncontrollable and is often ignored as a black box variable. Sometimes video quality problems are not even classified as events and cannot be incorporated into the analysis model, thus lacking a systematic response strategy.
[0113] The technical solution of the present invention for the first time takes the abnormality of video signal quality as part of the recognition event itself, and uses the lower limit of the acceptable bit rate and the upper limit of the packet loss rate as independent judgment thresholds. This logic can establish an input credibility model for the recognition system, linking data quality with the result of behavior judgment, similar to giving the system self-awareness, that is, judging whether the system can see clearly enough. Such a mechanism ensures the effectiveness of data from the input source and is a key link in the closed-loop of the overall quality control of the recognition system. It can not only significantly improve the fault tolerance ability of the system under low-quality video sources, but also automatically mark and eliminate unreliable recognition segments, improving the credibility of the overall event judgment. By introducing the dimension of signal source quality perception into the traffic video analysis system, the integrity of the system is improved, enabling the system to have the ability to monitor meta-events, that is, not only to recognize traffic events, but also to recognize technical failures that affect traffic event recognition. Timely perception and diagnosis of the root cause of model output distortion are carried out to construct a complete closed-loop of data-model-judgment link. In this way, once abnormal bit rate and packet loss are detected, the quality problem will be automatically reported, realizing the system's awareness of its own ability failure, avoiding blind decision-making or outputting false positive / false negative events, and avoiding the performance illusion during the cold start period from extending to the operation period.
[0114] The above-mentioned highway video cloud networking monitoring event analysis system runs on any computing device such as a desktop computer, a laptop computer, a palm computer or a cloud data center. The computing device includes: a processor, a memory, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it realizes the steps in the above-mentioned highway video cloud networking monitoring event analysis method. The operable system may include, but is not limited to, a processor, a memory, and a server cluster.
[0115] The embodiment of the present invention provides a highway video cloud networking monitoring event analysis system, as Figure 2 shown. The highway video cloud networking monitoring event analysis system of this embodiment includes: a processor, a memory, and a computer program stored in the memory and operable on the processor. When the processor executes the computer program, it realizes the steps in the above-mentioned embodiment of the highway video cloud networking monitoring event analysis method. When the processor executes the computer program, it runs in the following units of the system:
[0116] A data acquisition unit, configured to access the video stream collected by the roadside camera into the analysis system through the edge gateway, and attach data such as the camera number, geographical coordinates, and time stamp to each frame of the video stream. The analysis system performs analysis, analyzes and identifies the moving speed and angle of the vehicle in each frame, and analyzes and outputs the historical correct detection rate and the historical false alarm rate.
[0117] A variable operation unit, configured to dynamically generate vehicle flow evidence score components, stay frame thresholds, turning angle mutation evidence score components, accident discrimination duration thresholds, and restricted area stay duration thresholds based on data of video streams collected by roadside cameras;
[0118] A target recognition unit, configured to apply real-time target detection and tracking algorithms to each frame of image, assign a unique identifier to each pedestrian or vehicle, and calculate its instantaneous speed and the change in turning angle between adjacent frames;
[0119] A target judgment unit, configured to, for each tagged target, when its speed continuously drops below the vehicle flow evidence score component and the number of low-speed frames reaches or exceeds the stay frame threshold, determine that a traffic congestion event has occurred; when the change in turning angle of a single frame exceeds the turning angle mutation evidence score component and its speed remains below the vehicle flow evidence score component within the accident discrimination duration threshold subsequently, determine that an accident event has occurred; for the center of the bounding box of each tagged target, when the center of its bounding box enters a set prohibited area and the continuous stay time reaches or exceeds the restricted area stay duration threshold, determine that an illegal intrusion event has occurred;
[0120] A judgment transmission unit, configured to aggregate multiple discrimination results of the same identifier, the same event type, and adjacent or overlapping in time, and output them as a complete event record, where the event record includes the event type, target identifier, camera number to which it belongs, start and end times, and geographical coordinates.
[0121] Among them, in order to better unify the linear relationship and probability connection of the numerical values between physical quantities of different units, dimensionless processing can be performed between different physical quantities.
[0122] Among them, preferably, for all undefined variables in the present invention, if there is no clear definition, they can all be manually set thresholds.
[0123] The highway video cloud-connected monitoring event analysis system can run on computing devices such as desktop computers, laptop computers, palm computers, and cloud data centers. The highway video cloud-connected monitoring event analysis system includes, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above examples are only examples of the highway video cloud-connected monitoring event analysis method, system, and device, and do not constitute a limitation on the highway video cloud-connected monitoring event analysis method, system, and device. It can include more or fewer components than the examples, or combine certain components, or different components. For example, the highway video cloud-connected monitoring event analysis system can also include input / output devices, network access devices, buses, etc.
[0124] The present invention also provides an electronic device, a readable storage medium, and a computer program product:
[0125] An electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0126] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0127] A computer program product includes a computer program, and the computer program, when executed by a processor, implements the method for analyzing highway video cloud networking monitoring events and each step thereof.
[0128] Among them, the electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0129] Various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0131] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0134] A computer system can include a client and a server. The client and the server are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.
[0135] The so-called processor can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete component gate circuits, or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor is the control center of the highway video cloud networking monitoring event analysis system and uses various interfaces and lines to connect each sub-region of the entire highway video cloud networking monitoring event analysis system.
[0136] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and invoking the data stored in the memory, the processor realizes various functions of the method, system and device for analyzing events based on highway video cloud networking. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0137] It should be understood that various forms of the processes shown above can be used, and steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present invention can be achieved. No limitation is imposed herein.
[0138] The present invention relates to a method, system and device for event analysis based on highway video cloud networking. It can collect video streams through roadside cameras and, in combination with a pre-trained vision model, perform real-time identification and tracking of vehicle speed, steering angle, etc. The system dynamically generates traffic flow evidence score components, stay frame thresholds, corner mutation evidence score components, accident discrimination duration thresholds, and restricted area stay duration thresholds according to historical correct detection rates and false alarm rates, which are respectively used for the determination of congestion, accident and illegal intrusion events. A video bit rate and packet loss rate monitoring mechanism is also introduced to mark and report abnormal video quality. This method has an adaptive adjustment ability, is applicable to long-term operation scenarios after model cold start, and can effectively improve the accuracy of traffic event recognition and the robustness of the system. The present invention uses observable physical quantities and real-time statistical data to drive key parameters, and has high adaptability and scenario generality.
[0139] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Based on the analysis method of highway video cloud networking monitoring events, the video stream collected by roadside cameras is accessed to the analysis system through an edge gateway. The analysis system analyzes each frame of image, identifies target objects and calculates their moving speed and steering angle, and analyzes and outputs the historical correct detection rate and historical false alarm rate. It is characterized in that, The method includes: dynamically generating traffic flow evidence ratio components, stay frame thresholds, turning angle mutation evidence ratio components, accident discrimination duration thresholds, and restricted area stay duration thresholds based on data from video streams collected by roadside cameras: The traffic flow evidence ratio components are generated by combining adjusted evidence ratio factors with road speed limits and historical free flow speeds, where the historical detection evidence ratio is calculated using the historical correct detection rate and historical false alarm rate, and the adjusted evidence ratio factor is formed through the historical detection evidence ratio; the stay frame threshold is obtained by rounding up the product of the shortest duration required for congestion determination and the video sampling rate; the turning angle mutation evidence ratio component is obtained by adding the mean of the turning angle changes to the combination of the turning angle mutation baseline and the historical detection evidence ratio; the accident discrimination duration threshold is obtained by adding the congestion determination duration to a pre-set accident delay compensation duration; the restricted area stay duration threshold is obtained by dividing the physical length of the restricted area by the real-time average vehicle speed on the same road segment; applying object detection and tracking algorithms to each frame of the image, assigning a unique identifier to each target object, and calculating the speed and angle changes in real time; for each tagged target, obtaining its speed and the number of low-speed frames, and determining whether a traffic congestion event occurs by combining the traffic flow evidence ratio components and the stay frame threshold; obtaining the single-frame turning angle change and speed change of each tagged target, and determining whether an accident event occurs by combining the turning angle mutation evidence ratio components and the traffic flow evidence ratio components; when the center of the bounding box of each tagged target enters a set restricted area and the continuous stay time reaches or exceeds the restricted area stay duration threshold, determining that an illegal intrusion event occurs; aggregating multiple recognition results with the same identifier, the same event type, and adjacent or overlapping in time to form a complete event record.
2. The method according to claim 1, characterized in that, The analysis system uses a real-time detection network based on deep learning and a multi-target association tracking algorithm to achieve the positioning and recognition of each frame of the target and the real-time calculation and annotation of its speed and angle change amounts.
3. The method according to claim 1, characterized in that The stay frame threshold is dynamically calculated and rounded up based on the shortest congestion determination time and the actual sampling frame rate of the video stream; among them, the average flow rate is calculated based on the speeds of each vehicle in the data of the video stream collected by the roadside camera, and vehicles with speeds lower than 10% of the average flow rate in the data of the video stream are judged as congested vehicles, the average of the stay times of the congested vehicles is used as the shortest congestion determination time, and the frames with speeds lower than 10% of the average flow rate in the data of the video stream are used as low-speed frames, and the number of low-speed frames is used as the number of low-speed frames.
4. The method according to claim 3, wherein The accident discrimination duration threshold is obtained by adding the shortest congestion determination time and the accident delay compensation time, where the accident delay compensation time is the average of the times from when each vehicle in the video stream collected by the roadside camera stops moving to when it resumes moving.
5. The method according to claim 3, wherein The restricted area stay duration threshold is calculated based on the ratio of the physical length of the restricted area to the real-time average driving speed of this road segment, where the physical length of the restricted area is the length obtained by proportionally scaling or actually measuring the set restricted area in the physical space in the video stream.
6. The method according to claim 2, wherein Among them, The traffic flow certificate ratio component is generated by combining the adjusted certificate ratio factor with the road speed limit and the historical free flow speed, and specifically further includes: Obtain the historical correct detection rate, historical false alarm rate, target detection rate, and target false alarm rate; wherein, the historical correct detection rate is the frequency of detecting real congestion events within a period of time; the historical false alarm rate is the frequency of misjudging normal traffic as congestion within the same period of time; The ratio of the historical correct detection rate to the historical false alarm rate is the historical detection certificate ratio, the natural logarithm function of the historical detection certificate ratio is the historical empirical certificate ratio, and the ratio of the historical empirical certificate ratio to one hundred is the adjusted certificate ratio factor; Divide the historical false alarm rate by the historical detection certificate ratio to obtain the target false alarm rate, and the part obtained by subtracting the target false alarm rate from the whole one is the target detection rate; Select a sampling time window, obtain the correct detection rate and false alarm rate within the sampling time window. If the value of the correct detection rate is less than the target detection rate, then update the specific value of the adjusted certificate ratio factor by multiplying the specific value of the adjusted certificate ratio factor by the ratio of the historical correct detection rate divided by the target detection rate; Obtain the road speed limit of the section to be detected, and obtain the arithmetic mean of the vehicle flow speeds in the past period of time (which can be multiple different sampling time windows in the past) of the section to be detected as the historical free flow speed, and combine the minimum value of the obtained road speed limit and the historical free flow speed with the updated adjusted certificate ratio factor to generate a driving speed threshold.
7. The method according to claim 6, wherein Wherein, The turning angle mutation certificate ratio component is obtained by adding the mean value of the statistical turning angle change to the combination of the historical detection certificate ratio and the turning angle mutation benchmark, and specifically further includes: Obtain the maximum value in the turning angle difference sequence within the sampling time window as the turning angle difference peak value; Calculate the ratio of the standard deviation of the turning angle sequence to the mean value of the turning angle sequence as the sequence turning angle coordinate, and multiply the sequence turning angle coordinate by the turning angle difference peak value to obtain the turning angle mutation benchmark. The mean value of the turning angle sequence plus the combination of the historical detection certificate ratio and the turning angle mutation benchmark obtains the turning angle mutation certificate ratio component; Wherein, the historical detection certificate ratio and the turning angle mutation benchmark are combined in a multiplicative manner.
8. The method according to claim 6, characterized in that It also includes that in the stage where the video stream collected by the roadside camera is accessed to the analysis system through the edge gateway, the average bit rate and network packet loss rate of the video stream are monitored in real time to obtain the real-time packet loss rate, and the video quality is marked as abnormal and then uploaded as a special event through the acceptable bit rate certificate ratio lower limit and the packet loss rate certificate ratio upper limit; wherein, the positive reporting rate in multiple different sampling time windows is continuously monitored and the mean value of the positive reporting rate is obtained. The acceptable bit rate certificate ratio lower limit is automatically taken as the product of the average bit rate and the mean value of the positive reporting rate according to historical statistics, the packet loss rate certificate ratio upper limit is the product of the average packet loss rate and the historical detection certificate ratio, and the average packet loss rate is the arithmetic mean of the real-time packet loss rate.
9. Highway video cloud networking monitoring event analysis system, characterized in that, The above-mentioned highway video cloud networking monitoring event analysis system runs on any computing device of a desktop computer, a laptop computer or a cloud data center. The computing device includes: a processor, a memory, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps in the highway video cloud networking monitoring event analysis method according to any one of claims 1 to 8.
10. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Intelligent traffic monitoring system and monitoring method
CN114738627B
An on-board monitoring system for intelligent traffic inspection vehicles
CN118247761B
Traffic incident monitoring method and system
CN118865278A
Road state monitoring system based on video image analysis
CN119516765A
Event detection method and apparatus for cloud control platform, device, and storage medium
US20210224553A1