Traffic data analysis method and device based on multi-channel video fusion

By using multi-channel video fusion technology, multiple video streams from traffic monitoring systems are acquired and processed, solving the problem of low data processing efficiency and enabling efficient and accurate analysis of traffic data and real-time identification of abnormal targets.

CN121033780APending Publication Date: 2025-11-28JIANGSU YIJIESI INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511160146.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing traffic monitoring systems suffer from low data processing efficiency due to the lack of an efficient multi-source data fusion mechanism, especially in real-time monitoring scenarios where high-volume video streams can cause processing delays.

Method used

By using multi-channel video fusion technology, a set of monitoring nodes in the target traffic area is acquired, standardized and downsampled, background is identified and foreground objects are extracted, temporal arrangement and spatial coding are performed, panoramic video sequence frames are generated, and traffic operation parameters and abnormal target information are tracked and processed in parallel.

Benefits of technology

It has improved the accuracy and comprehensiveness of traffic data analysis, enabled real-time monitoring of traffic conditions and automatic identification of abnormal behaviors, and reduced the occurrence of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033780A_ABST
    Figure CN121033780A_ABST
Patent Text Reader

Abstract

The invention provides a traffic data analysis method and device based on multi-channel video fusion, and relates to the technical field of traffic condition analysis, and the method comprises the steps: collecting a multi-channel traffic video stream set; standardization processing and down-sampling processing are carried out, and background identification removal is carried out on the multi-path traffic video frame set; arranging the multi-path traffic foreground image frame set according to time sequence information, carrying out traffic data analysis and abnormal target identification, and determining multi-path traffic operation parameters and abnormal traffic target information; space coding is carried out, and alignment fusion is carried out on the multiple traffic foreground sequence frame sets; and performing parallel tracking processing to obtain a target traffic data analysis result. Through the method and the device, the technical problem of low data processing efficiency caused by lack of an efficient multi-source data fusion mechanism in the prior art can be solved, and the accuracy and comprehensiveness of traffic data analysis are improved through a multi-channel video fusion technology and selection and deployment of monitoring nodes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of traffic condition analysis, and particularly relates to a traffic data analysis method and device based on multi-path video fusion. BACKGROUND

[0002] Traffic management and monitoring is an important part of urban management. Traditional traffic monitoring systems usually rely on a single camera or a limited number of cameras, which has obvious limitations in dealing with complex traffic conditions. A single camera cannot provide a comprehensive traffic view, and there are problems of data synchronization and low processing efficiency between multiple cameras. Secondly, existing technologies cannot effectively process large-scale video data streams, especially in real-time monitoring scenarios, high data volume of video streams may cause processing delay. Limited by hardware resources such as processor computing power, memory capacity, etc., the efficiency of data processing is limited.

[0003] In summary, the existing technology has the technical problem of low data processing efficiency due to the lack of an efficient multi-source data fusion mechanism. SUMMARY

[0004] The purpose of the present application is to provide a traffic data analysis method and device based on multi-path video fusion, to solve the technical problem of low data processing efficiency due to the lack of an efficient multi-source data fusion mechanism in the prior art.

[0005] In view of the above problems, the present application provides a traffic data analysis method and device based on multi-path video fusion.

[0006] In a first aspect, the application provides a traffic data analysis method based on multi-channel video fusion, which is implemented by a traffic data analysis device based on multi-channel video fusion. The traffic data analysis method based on multi-channel video fusion comprises: acquiring a target traffic area, deploying monitoring nodes on the target traffic area, obtaining a set of traffic video monitoring nodes, and collecting a set of multi-channel traffic video streams through the set of traffic video monitoring nodes; performing standardization processing and downsampling processing on the set of multi-channel traffic video streams respectively, obtaining a set of multi-channel traffic video frames, and performing background recognition and removal on the set of multi-channel traffic video frames to obtain a set of multi-channel traffic foreground image frames; arranging the set of multi-channel traffic foreground image frames according to time sequence information to obtain a set of multi-channel traffic foreground sequence frames, and performing traffic data analysis and abnormal target identification on the set of multi-channel traffic foreground sequence frames to determine multi-channel traffic operation parameters and abnormal traffic target information; performing spatial coding on the set of multi-channel traffic foreground sequence frames to obtain a set of multi-channel traffic video identification codes, aligning and fusing the set of multi-channel traffic foreground sequence frames according to the set of multi-channel traffic video identification codes to generate a traffic area panoramic video sequence frame; and performing parallel tracking processing on the multi-channel traffic operation parameters and abnormal traffic target information based on the traffic area panoramic video sequence frame to obtain a target traffic data analysis result.

[0007] In a second aspect, the application further provides a traffic data analysis device based on multi-channel video fusion, configured to execute the traffic data analysis method based on multi-channel video fusion as described in the first aspect. The traffic data analysis device based on multi-channel video fusion comprises: a data acquisition module configured to acquire a target traffic area, deploy monitoring nodes in the target traffic area, obtain a set of traffic video monitoring nodes, and collect a set of multi-channel traffic video streams through the set of traffic video monitoring nodes; a video preprocessing module configured to respectively perform standardization processing and down-sampling processing on the set of multi-channel traffic video streams, obtain a set of multi-channel traffic video frames, and perform background recognition and removal on the set of multi-channel traffic video frames to obtain a set of multi-channel traffic foreground image frames; a feature extraction module configured to arrange the set of multi-channel traffic foreground image frames according to time sequence information to obtain a set of multi-channel traffic foreground sequence frames, and perform traffic data analysis and abnormal target recognition on the set of multi-channel traffic foreground sequence frames to determine multi-channel traffic operation parameters and abnormal traffic target information; a data fusion module configured to perform spatial coding on the set of multi-channel traffic foreground sequence frames to obtain a set of multi-channel traffic video identification codes, align and fuse the set of multi-channel traffic foreground sequence frames according to the set of multi-channel traffic video identification codes to generate traffic area panoramic video sequence frames; and a parallel tracking processing module configured to perform parallel tracking processing on the multi-channel traffic operation parameters and abnormal traffic target information based on the traffic area panoramic video sequence frames to obtain target traffic data analysis results.

[0008] One or more technical solutions provided in the application have at least the following technical effects or advantages: By acquiring a target traffic area, a monitoring node is deployed in the target traffic area to obtain a traffic video monitoring node set, and a plurality of traffic video stream sets are collected through the traffic video monitoring node set; the plurality of traffic video stream sets are respectively subjected to standardization processing and down-sampling processing to obtain a plurality of traffic video frame sets, and the plurality of traffic video frame sets are subjected to background recognition and removal to obtain a plurality of traffic foreground image frame sets; the plurality of traffic foreground image frame sets are arranged according to time sequence information to obtain a plurality of traffic foreground sequence frame sets, and the plurality of traffic foreground sequence frame sets are subjected to traffic data analysis and abnormal target recognition to determine a plurality of traffic running parameters and abnormal traffic target information; the plurality of traffic foreground sequence frame sets are subjected to spatial coding to obtain a plurality of traffic video identification code sets, and the plurality of traffic foreground sequence frame sets are aligned and fused according to the plurality of traffic video identification code sets to generate a traffic area panoramic video sequence frame; the plurality of traffic running parameters and abnormal traffic target information are subjected to parallel tracking processing based on the traffic area panoramic video sequence frame to obtain a target traffic data analysis result. That is, through the multi-channel video fusion technology, the monitoring nodes are selected and deployed, and the accuracy and comprehensiveness of traffic data analysis are improved.

[0009] The above description is only a summary of the technical solutions of the present application. In order to enable one skilled in the art to better understand the technical means of the present application, the present application can be implemented according to the content of the specification, and in order to enable the above and other purposes, features and advantages of the present application to be more obvious and easy to understand, the following specific embodiments of the present application are described. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only exemplary, and other drawings can be obtained by those skilled in the art without creating labor on the basis of the provided drawings.

[0011] Figure 1 Flowchart of the traffic data analysis method based on multi-channel video fusion of the present application; Figure 2 Structure diagram of the traffic data analysis device based on multi-channel video fusion of the present application.

[0012] Explanation of reference signs: data acquisition module 11, video preprocessing module 12, feature extraction module 13, data fusion module 14, and parallel tracking processing module 15. DETAILED DESCRIPTION

[0013] The application provides a traffic data analysis method and device based on multi-channel video fusion, which solves the technical problem of low data processing efficiency caused by the lack of an efficient multi-source data fusion mechanism in the prior art. The multi-channel video fusion technology is used to select and deploy monitoring nodes, thereby improving the accuracy and comprehensiveness of traffic data analysis.

[0014] The technical solutions in the application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the application, rather than all the embodiments of the application. It should be understood that the application is not limited to the example embodiments described herein. Based on the embodiments of the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the scope of protection of the application. In addition, it should be noted that, for the convenience of description, only parts related to the application are shown in the drawings rather than all parts.

[0015] Embodiment one, please refer to the accompanying Figure 1 The application provides a traffic data analysis method based on multi-channel video fusion, which is applied to a traffic data analysis device based on multi-channel video fusion. The traffic data analysis method based on multi-channel video fusion specifically includes the following steps. Step one: Obtain a target traffic area, deploy monitoring nodes in the target traffic area, obtain a traffic video monitoring node set, and collect a multi-channel traffic video stream set through the traffic video monitoring node set.

[0016] Specifically, the traffic area to be monitored is determined, which is usually a place with large traffic flow, frequent accidents, or high traffic management demand. For example, a busy intersection in a city, a section of a highway, or a traffic hub area. In the determined target traffic area, multiple high-definition cameras are deployed as monitoring nodes, which are usually located at key traffic nodes such as intersections, viaducts, and tunnels to ensure full coverage of the traffic area. The monitoring nodes work cooperatively to form a detection node set, collect video streams from multiple perspectives, and form a multi-channel traffic video stream set. Each camera provides a unique perspective, which together constitutes a comprehensive monitoring of the entire traffic area. By deploying cameras at multiple key nodes, the traffic area can be fully covered, and the comprehensiveness of monitoring can be improved.

[0017] Step two: Standardize and downsample the multi-channel traffic video stream set respectively to obtain a multi-channel traffic video frame set, and remove the background of the multi-channel traffic video frame set to obtain a multi-channel traffic foreground image frame set.

[0018] Specifically, the multi-channel traffic video stream set is standardized, including format standardization, filtering denoising, video compression and color space conversion, to ensure consistency in format, quality, data volume and color representation. The historical traffic flow data of the monitoring nodes of each camera is evaluated to obtain video traffic flow characteristic information, including traffic peak period, flow fluctuation mode, etc., to determine the appropriate video sampling step. If the traffic flow of a node video monitoring time is large, a smaller sampling step is used to collect more video frame data to obtain a multi-channel traffic video frame set, ensuring the detail and accuracy of data processing. Representative image frames with more background and foreground changes are selected from the multi-channel traffic video frame set to construct the initial background model of each video stream. Factors affecting the appearance of the background, such as light changes and weather changes, are identified. The Beijing model is updated according to these factors to adapt to dynamic background changes. The updated background model is used to identify the initial video frame set, separating moving foreground objects such as vehicles and pedestrians. Combining the multi-channel background update model and the multi-channel foreground object model, an attention recognition module is constructed to improve the accuracy and efficiency of foreground object recognition. The attention recognition module is used to process the video frame set, remove the background and retain the foreground objects, thereby obtaining a multi-channel traffic foreground image frame set. Through standardized processing, the quality and consistency of the video data are improved, and accurate background recognition removal helps to improve the accuracy and efficiency of subsequent data analysis, such as target detection, tracking and classification.

[0019] Step three: arrange the multi-channel traffic foreground image frame set according to the time sequence information to obtain a multi-channel traffic foreground sequence frame set, and perform traffic data analysis and abnormal target recognition on the multi-channel traffic foreground sequence frame set to determine the multi-channel traffic operation parameters and abnormal traffic target information.

[0020] Specifically, the multi-lane traffic foreground image frame set is arranged in time sequence according to the time when it is collected, and each video stream is converted into a set of foreground image frames arranged in time sequence, which constitutes a multi-lane traffic foreground sequence frame set. The arranged multi-lane traffic foreground sequence frame set is subjected to traffic data analysis, calculation and analysis of traffic operation parameters such as average speed, traffic flow, road congestion index, etc. An abnormal behavior standard is defined, and an abnormality recognition model is trained by machine learning algorithms such as deep learning, support vector machine, etc., so that it can recognize abnormal behaviors in the video. The multi-lane traffic foreground sequence frame set is processed using the abnormality recognition model to identify possible abnormal traffic targets such as speeding vehicles, vehicles driving in reverse, etc. Through traffic data analysis and abnormal target recognition, the running parameters and abnormal traffic target information of multi-lane traffic are determined. For example, the average speed, traffic flow, and the number and time points of abnormal behaviors such as speeding and driving in reverse at a crossroads within a day are determined. Through time sequence arrangement and analysis, the dynamic changes of traffic conditions over time are comprehensively understood, and accurate traffic data analysis and abnormal target recognition help to understand traffic patterns, improve traffic safety, and reduce traffic accidents.

[0021] Step four: spatially encoding the multi-lane traffic foreground sequence frame set to obtain a multi-lane traffic video identification code set, aligning and fusing the multi-lane traffic foreground sequence frame set according to the multi-lane traffic video identification code set to generate traffic area panoramic video sequence frames.

[0022] Specifically, the multi-lane traffic foreground sequence frame set is encoded according to the spatial position relationship of the monitoring nodes, taking into account the positions and angles of view of different cameras, ensuring the relevance and consistency of the encoding. Spatial encoding is the process of encoding each video stream according to the spatial position relationship of the monitoring nodes. Through spatial encoding, each video frame is converted into an identification code containing monitoring node position information, and these identification codes constitute a multi-lane traffic video identification code set. According to the identification code, the spatial parameters of each camera are extracted. The transformation matrix of each camera is calculated to convert the video frames from their original perspective to a unified world coordinate system. Using these transformation matrices, the multi-lane traffic foreground sequence frame set is mapped and fused to obtain panoramic traffic fusion video information. According to the timestamp information of each video frame, the frame order in the panoramic video is adjusted, and the video frames of different cameras are aligned and fused to ensure that the scenes captured by all cameras at the same time are aligned in time, thereby generating traffic area panoramic video sequence frames. Through spatial encoding and the use of identification code sets, unified identification and association of multi-lane video data are achieved, improving the efficiency and accuracy of data processing.

[0023] Step five: parallel tracking processing of the multi-channel traffic operation parameters and abnormal traffic target information based on the traffic area panoramic video sequence frames, to obtain target traffic data analysis results.

[0024] Specifically, the traffic operation parameters of key road nodes in the traffic panoramic area are extracted and calculated through the traffic area panoramic video sequence frames. Abnormal traffic target information in the video, such as speeding, reverse driving, and red light running, is extracted. A feature tracking network is constructed to track the target trajectory of the panoramic video frames according to the appearance feature data of each abnormal target, such as pedestrians or cars, to obtain the running trajectory of each abnormal target, which records the moving path and behavior pattern of the abnormal target in the video sequence. The target traffic data analysis results are obtained by integrating the key road node traffic operation parameter information and the abnormal target behavior feature parameters, including comprehensive evaluation of traffic conditions and detailed analysis of abnormal traffic behavior. Through parallel tracking processing, real-time monitoring and analysis of traffic conditions can improve the efficiency and accuracy of traffic monitoring.

[0025] Further, step two of the present application includes: A video standardization process is obtained, which includes format standardization, filter denoising, video compression, and color space conversion. The multi-channel traffic video stream set is standardized based on the video standardization process to obtain a multi-channel standard traffic video stream set. The traffic video monitoring node set is analyzed for traffic characteristics to obtain video traffic characteristic information, and a multi-channel video sampling step set is determined based on the video traffic characteristic information. The multi-channel standard traffic video stream set is down-sampled based on the multi-channel video sampling step set to obtain the multi-channel traffic video frame set.

[0026] Specifically, the set of multi-channel traffic video streams is standardized, and each video stream will go through format standardization, filtering denoising, video compression, and color space conversion processes to ensure consistency in format, quality, compression rate, and color space. Specifically, different formats of video files are converted to a unified file format such as MP4, AVI, MKV, etc. through format standardization. For example, MOV, FLV, and other formats are converted to more general or more suitable formats for subsequent processing, such as MP4, AVI, or MKV. Various filtering algorithms such as mean filtering, median filtering, or Gaussian filtering are used to remove noise in the video and improve video quality. In order to save storage space and improve transmission efficiency, videos often need to be compressed. Selecting the appropriate compression algorithm can maintain video quality while reducing file size for more efficient storage and transmission. Common video compression standards include MPEG, H.264, H.265, etc. Different video sources may use different color spaces (such as RGB, YUV, etc.). Color space conversion is the process of converting video data from one color space to another, aiming to convert all videos to a unified color space. Common color spaces include RGB, YUV, HSV, etc.

[0027] The monitoring nodes of each camera are evaluated based on historical traffic flow data, analyzing the traffic flow of each node at different time periods, and video monitoring time and other characteristic parameters. Based on the results of the flow characteristic analysis, the video sampling step of each monitoring node is determined. If the traffic flow of the monitoring node is more at the video monitoring time, the sampling step is smaller, meaning more video frame data is collected. For busy intersections, the sampling step may be set to 30 frames per second to ensure that all traffic details are captured. For road sections with less traffic flow, the sampling step is set to 15 frames per second, which is sufficient to capture the necessary traffic information while reducing the burden of data processing and saving storage space. Based on the above-determined sampling step, the set of multi-channel standard traffic video streams is downsampled. Downsampling refers to the process of extracting frames from a video stream based on a pre-determined sampling step. The sampling step determines the frequency of extracting video frames. After downsampling, each video stream is converted into a set of video frames, which form a set of multi-channel traffic video frames. Through standardized processing, all video streams are converted to a unified format and standard quality, and the multi-channel video stream is standardized and downsampled to ensure video quality while improving processing efficiency and compatibility.

[0028] Further, the present application further comprises the following steps: The initial frame selection is performed on the multi-channel traffic video frame set to obtain an initial multi-channel video frame set, and initial multi-channel background models are constructed based on the initial multi-channel video frame set. Background transformation factors are obtained, and the initial multi-channel background models are updated based on the background transformation factors to obtain multi-channel background update models. Based on the multi-channel background update models, foreground object recognition is performed on the initial multi-channel video frame set to obtain multi-channel foreground object models. Based on the multi-channel background update models and the multi-channel foreground object models, an attention recognition module is constructed, and background recognition removal is performed on the multi-channel traffic video frame set based on the attention recognition module to obtain the multi-channel traffic foreground image frame set.

[0029] Specifically, initial image frames in the multi-channel traffic video frame set are selected. The selection of initial frames should take into account the changes in background and foreground. The ideal choice is those frames with stable background and rich foreground changes, because such frames can provide enough information to build a comprehensive background model. Representative background usually refers to frames that should contain complete background information of the scene, such as roads, buildings, traffic signs, etc. Based on the selected initial frames, a background model is constructed for each video stream through background modeling algorithms such as mean value method, median method or Gaussian mixture model. Background model is an important tool for distinguishing background and foreground (such as vehicles and pedestrians) in video frames.

[0030] The transformation factors affecting the background model are identified and obtained, including illumination changes, weather changes, seasonal changes, etc. Based on the background transformation factors, the initial multi-channel background models are updated. The purpose of updating is to enable the background model to adapt to environmental changes, so as to more accurately separate foreground and background. The update strategy can be based on the temporal stability of pixels, i.e. only updating the pixels that remain stable for a long time, or using more complex algorithms such as Gaussian mixture model. Gaussian mixture model is a statistical model that represents multiple patterns in the background, thereby better adapting to dynamic background changes. The initial multi-channel video frame set is subjected to foreground object recognition using the multi-channel background update model, and image processing techniques such as frame difference, background subtraction, etc. are used to identify and separate foreground objects. Foreground object recognition is the process of distinguishing foreground (such as vehicles and pedestrians) and background in video frames. Through foreground object recognition, foreground objects can be extracted from each video stream to form multi-channel foreground object models, representing moving objects in different video streams.

[0031] According to the multi-path background update model and the multi-path foreground object model, an attention recognition module is constructed. The attention recognition module is a high-level image processing technique that focuses on key areas in video frames, such as moving objects or areas of interest, while ignoring irrelevant background information. Through the constructed attention recognition module, the multi-path traffic video frame set is subjected to background recognition removal, and the foreground (such as vehicles and pedestrians) is separated from the video frames, thereby obtaining a multi-path traffic foreground image frame set. After background recognition removal, each video stream is converted into a set of image frames containing only foreground objects, which form the multi-path traffic foreground image frame set. By selecting representative initial frames and dynamically updating the background model, the foreground objects are accurately recognized and segmented, improving the accuracy of traffic data analysis.

[0032] Further, step three of the present application includes: A set of traffic feature indicators is obtained, and based on the set of traffic feature indicators, the set of multi-path traffic foreground sequence frames is subjected to associated feature extraction, obtaining a set of multi-path traffic indicator associated features. The set of multi-path traffic indicator associated features is subjected to indicator calculation according to the set of traffic feature indicators, and the multi-path traffic running parameters are determined. The abnormal behavior of each foreground object in the multi-path foreground object model is defined, and the abnormal behavior standard rules of the foreground objects are obtained. Based on the abnormal behavior standard rules of the foreground objects, data crawling classification and abnormal target identification training are performed, and a foreground target abnormality recognition model is constructed. Based on the foreground target abnormality recognition model, the set of multi-path traffic foreground sequence frames is subjected to abnormal target recognition, and the abnormal traffic target information is determined.

[0033] Specifically, a set of traffic feature indicators is obtained, including the number of vehicles, speed, lane occupancy rate, traffic density, etc. Image processing techniques (such as edge detection, corner detection, etc.) and convolutional neural networks (CNN) are used to extract associated traffic indicator features from each image. For example, the position and speed of a vehicle are detected using image processing techniques, thereby extracting the feature of vehicle speed. Through associated feature extraction, traffic indicator features are extracted from each video stream, forming a set of multi-path traffic indicator associated features. Using the extracted features, the set of multi-path traffic indicator associated features is calculated to determine the running parameters of the multi-path traffic, such as average speed, traffic flow, road congestion index, etc. For example, the number of vehicles passing a certain point within a certain time period is calculated to obtain the traffic flow. For each foreground object (such as vehicles, pedestrians, etc.) in the multi-path foreground object model, the definition of abnormal behavior is set, including the definition of speeding, reverse driving, and red light running traffic violations. The definition of speeding is that the speed of a vehicle exceeds the road speed limit standard, the definition of reverse driving is that a vehicle drives in a reverse lane, and the definition of red light running is that a vehicle passes through an intersection when the light is red. According to the defined abnormal behavior, standard rules are formulated for judging and identifying abnormal behavior.

[0034] Using predefined abnormal behavior criteria rules, foreground objects in multiple traffic video frames are classified. Normal traffic behaviors and abnormal behaviors (such as speeding, driving against traffic, running red lights, etc.) are identified and classified. Simultaneously, abnormal behavior labels are trained, and machine learning algorithms (such as deep learning) are used to identify and distinguish between normal and abnormal behaviors. Through training, a model capable of identifying and classifying abnormal behaviors of foreground targets is built, automatically recognizing abnormal behaviors such as speeding, driving against traffic, and running red lights. The foreground target anomaly recognition model is used to process a collection of foreground sequence frames from multiple traffic sources to identify abnormal traffic targets. By analyzing each frame image through the model, abnormal traffic targets that violate traffic rules (such as speeding vehicles, vehicles driving against traffic, etc.) are identified, and specific information about the abnormal traffic targets, such as their location, type, and behavior, is determined.

[0035] In a specific example, assume the acquired traffic feature indicators include the number of vehicles, average speed, and lane occupancy rate. The feature extraction results from 1000 frames of video data are: an average of 50 vehicles per frame, an average speed of 30 km / h, and a lane occupancy rate of 70%. The calculated average speed over 10 hours is 25 km / h, the traffic flow is 5000 vehicles / hour, and the road congestion index is 3 (on a scale of 1-5, where 5 represents congestion). The defined abnormal behavior rules are: exceeding 50 km / h is considered speeding, driving against traffic, and running a red light. Using a deep learning model, 5000 normal frames and 500 abnormal frames were trained, achieving a 98% accuracy rate for normal frame recognition and a 95% accuracy rate for abnormal frame recognition. Processing 10000 frames of video data, 50 abnormal targets were identified, including: 20 speeding vehicles, 15 vehicles driving against traffic, and 15 vehicles running red lights. By extracting associated features and calculating indicators, accurate traffic operation parameters are obtained. The construction of abnormal behavior definition and identification models enables the system to automatically detect and report abnormal traffic events, effectively monitor and analyze traffic conditions at intersections, and improve the level of intelligence in traffic monitoring.

[0036] Furthermore, step four of this application includes: Based on the set of multi-channel traffic video identifiers, spatial attribute information and temporal attribute information of multi-channel video monitoring are determined; a calibration video perspective is selected, and the spatial attribute information of the multi-channel video monitoring is transformed and analyzed according to the calibration video perspective to obtain a set of multi-channel video data transformation matrices; based on the set of multi-channel video data transformation matrices, the set of multi-channel traffic foreground sequence frames is mapped, transformed, and fused to obtain panoramic traffic fused video information; based on the temporal attribute information of the multi-channel video monitoring, the image frames in the panoramic traffic fused video information are time-synchronized and aligned to generate the panoramic video sequence frames of the traffic area.

[0037] Specifically, according to the multi-channel traffic video identification code set, the spatial attribute information of each video frame is extracted, including the position of the camera in three-dimensional space, the orientation and angle of the camera, the resolution, etc. According to the multi-channel traffic video identification code set, the time attribute information of each video frame is extracted, including the acquisition time of the frame, the timestamp, etc., which records the exact time when the video frame is captured. A reference camera angle is selected as the calibration angle, which usually has a stable field of view and good coverage. According to the spatial parameters of the camera (including position, orientation, etc.), the spatial attribute information of the multi-channel video monitoring is converted and analyzed, and the foreground frame of each channel video is transformed from the original angle to the selected calibration video angle. Usually, spatial transformation such as rotation matrix, affine transformation matrix, etc. is needed. Through conversion and analysis, a video data transformation matrix is generated for each camera, forming a multi-channel video data transformation matrix set.

[0038] Using these transformation matrices, each video frame in the multi-channel traffic foreground sequence frame set is spatially transformed according to its corresponding transformation matrix, realizing the unification and fusion of different video angles. Through mapping and transformation fusion, a panoramic traffic fusion video is obtained, which integrates information from different cameras and provides a comprehensive and continuous traffic scene view. Using the extracted multi-channel video monitoring time attribute information, the image frames in the panoramic traffic fusion video information are time-synchronized and aligned. The timestamps of different camera video frames are compared, and the time order is adjusted to ensure consistency in time. The multi-channel video image frames that have been time-synchronized and aligned are merged to generate a time-synchronized panoramic traffic fusion video sequence frame. This sequence frame contains video data from multiple monitoring nodes and is completely time-aligned, providing a continuous panoramic view video. Through accurate spatial and temporal attribute information determination, efficient angle conversion and transformation matrix, the unified angle display and time synchronization of multi-channel video data are realized, providing a comprehensive and continuous traffic scene view.

[0039] Further, step five of the present application includes: Based on the traffic area panoramic video sequence frame, the multi-channel traffic running parameter is calculated for key road node fitting, the key road node traffic running parameter information is determined, the anchor frame identification and feature extraction are performed on the abnormal traffic target information in turn, and a plurality of abnormal target feature data sets are obtained. Based on the plurality of abnormal target feature data sets, a feature tracking network is built, the plurality of abnormal target feature data sets and the traffic area panoramic video sequence frame are tracked and analyzed in parallel by using the feature tracking network, a plurality of abnormal target behavior feature parameters are obtained, and the target traffic data analysis result is obtained based on the key road node traffic running parameter information and the plurality of abnormal target behavior feature parameters.

[0040] Specifically, the traffic area panoramic video sequence frames are analyzed to identify key road nodes, including intersections with heavy traffic, junctions of major roads, etc. Traffic operation parameters are extracted and calculated for the key road nodes. If a main road is connected by multiple auxiliary roads, the traffic parameters of the main road can be calculated based on the traffic parameters of the auxiliary roads. Anchor boxes are identified for abnormal traffic targets (such as pedestrians, illegal vehicles, etc.) in the video sequence frames, and an anchor box is generated for each target, which marks the position and size of the target in the image. Key features are extracted from the anchor box, including the shape, size, color, texture, etc. of the target. The features extracted from multiple abnormal targets are organized into data sets, which contain the feature information of abnormal targets.

[0041] The multiple abnormal target feature data sets contain features extracted from different abnormal targets, such as the appearance features of pedestrians or cars. A network capable of tracking target trajectories based on feature data is built using deep learning techniques, such as a twin network architecture. Using the built feature tracking network, the positions and motion trajectories of specific targets in the video frames are identified and tracked based on the appearance feature data of each abnormal target. The trained feature tracking network is applied to the panoramic video sequence frames to track abnormal targets in real time. For each frame in the video sequence, the detected abnormal targets are identified and labeled, and the target trajectory information is updated. The running trajectory of each abnormal target is associated with its behavior in the panoramic video sequence frames to obtain a behavior feature set for multiple abnormal targets, which contains behavior information related to each target trajectory.

[0042] Using the information in the traffic target abnormal behavior library, the behavior feature sets of multiple abnormal targets are matched and detected to determine the behavior feature parameters of each abnormal target, obtaining the behavior feature parameters of each abnormal target such as starting position, moving path, speed, acceleration, etc. By comprehensively analyzing the traffic operation parameters of key road nodes and the behavior features of abnormal targets, traffic data analysis results are obtained, including the causes of traffic congestion, high-risk areas of accidents, frequent locations of traffic rule violations, etc. The target traffic data analysis results refer to the conclusions or reports obtained after data analysis based on key road node traffic operation parameter information and abnormal target behavior feature parameters. By comprehensively analyzing the traffic situation and identifying traffic problems, combined with the analysis of traffic operation parameters and abnormal target behavior features, it helps to deeply understand the traffic pattern and improve the efficiency and accuracy of traffic monitoring.

[0043] Further, the present application further comprises the following steps: Parallel tracking and labeling of the multiple abnormal target feature data sets and the traffic area panoramic video sequence frames is performed using the feature tracking network, and multiple abnormal target running trajectories are obtained; a traffic target abnormal behavior library is constructed based on the foreground object abnormal behavior standard rules; target behavior correlation of the traffic area panoramic video sequence frames is performed according to the multiple abnormal target running trajectories, and multiple abnormal target behavior feature sets are obtained; behavior matching detection of the multiple abnormal target behavior feature sets is respectively performed based on the traffic target abnormal behavior library, and the multiple abnormal target behavior feature parameters are obtained.

[0044] Specifically, the feature tracking network is used to perform parallel tracking and labeling of the multiple abnormal target feature data sets and the traffic area panoramic video sequence frames, simultaneously track multiple abnormal targets, and label the positions and movements of the targets in the video frames while updating the trajectory information of the targets. Through parallel tracking and labeling, complete running trajectories of the multiple abnormal targets are generated, including the starting positions, moving paths, and final positions of each abnormal target. According to the foreground object abnormal behavior standard rules, different types of abnormal traffic behavior data are collected and sorted. The collected abnormal behavior data are organized into a database, each abnormal behavior has corresponding labels and descriptions, and the database is used to store and query related information of the abnormal behaviors.

[0045] According to the running trajectories of the abnormal targets, target behavior correlation of the traffic area panoramic video sequence frames is performed. The running trajectory of each abnormal target is associated with its behavior in the video frames, so that the behavior features of the multiple abnormal targets are obtained. The behavior feature set contains the behavior patterns and features of each abnormal target, such as driving speed, driving direction, etc. Through correlation analysis, the behavior feature set of the multiple abnormal targets is obtained, which contains the behavior information related to each target trajectory. The abnormal behavior library contains the definitions and discrimination criteria of various abnormal traffic behaviors, such as overspeed, reverse driving, and red light running. The behavior feature set of the multiple abnormal targets is matched and detected using the information in the behavior library. The trajectory data of each abnormal target is compared with the data in the behavior library to determine its behavior type and feature parameters. Through behavior matching detection, the behavior feature parameters of each abnormal target are determined. Through the feature tracking network and behavior matching detection, the abnormal behaviors in the video are automatically identified, and the intelligent level of the monitoring system is improved.

[0046] In summary, the traffic data analysis method based on multi-channel video fusion provided in the present application has the following technical effects: The target traffic area is acquired, monitoring nodes are deployed on the target traffic area, a traffic video monitoring node set is obtained, and a plurality of traffic video stream sets are collected through the traffic video monitoring node set; the plurality of traffic video stream sets are respectively subjected to standardization processing and down-sampling processing to obtain a plurality of traffic video frame sets, and the plurality of traffic video frame sets are subjected to background recognition removal to obtain a plurality of traffic foreground image frame sets; the plurality of traffic foreground image frame sets are arranged according to time sequence information to obtain a plurality of traffic foreground sequence frame sets, and the plurality of traffic foreground sequence frame sets are subjected to traffic data analysis and abnormal target recognition to determine a plurality of traffic running parameters and abnormal traffic target information; the plurality of traffic foreground sequence frame sets are subjected to spatial coding to obtain a plurality of traffic video identification code sets, and the plurality of traffic foreground sequence frame sets are aligned and fused according to the plurality of traffic video identification code sets to generate a traffic area panoramic video sequence frame; and the plurality of traffic running parameters and the abnormal traffic target information are subjected to parallel tracking processing based on the traffic area panoramic video sequence frame to obtain a target traffic data analysis result. That is, through the multi-channel video fusion technology, the monitoring nodes are selected and deployed, and the accuracy and comprehensiveness of traffic data analysis are improved.

[0047] In the second embodiment, based on the same inventive concept as the traffic data analysis method based on multi-channel video fusion in the first embodiment, the application also provides a traffic data analysis device based on multi-channel video fusion. Please refer to the accompanying drawings Figure 2 The traffic data analysis device based on multi-channel video fusion comprises: A data acquisition module 11 is configured to acquire a target traffic area, deploy monitoring nodes on the target traffic area, obtain a traffic video monitoring node set, and collect a plurality of traffic video stream sets through the traffic video monitoring node set.

[0048] A video preprocessing module 12 is configured to respectively subject the plurality of traffic video stream sets to standardization processing and down-sampling processing to obtain a plurality of traffic video frame sets, and subject the plurality of traffic video frame sets to background recognition removal to obtain a plurality of traffic foreground image frame sets.

[0049] A feature extraction module 13 is configured to arrange the plurality of traffic foreground image frame sets according to time sequence information to obtain a plurality of traffic foreground sequence frame sets, and subject the plurality of traffic foreground sequence frame sets to traffic data analysis and abnormal target recognition to determine a plurality of traffic running parameters and abnormal traffic target information.

[0050] The data fusion module 14 is configured to perform spatial coding on the multi-channel traffic foreground sequence frames to obtain a multi-channel traffic video identification code set, align and fuse the multi-channel traffic foreground sequence frames according to the multi-channel traffic video identification code set, and generate traffic region panoramic video sequence frames.

[0051] The parallel tracking processing module 15 is configured to perform parallel tracking processing on the multi-channel traffic running parameters and abnormal traffic target information based on the traffic region panoramic video sequence frames, and obtain target traffic data analysis results.

[0052] Further, the video preprocessing module 12 in the traffic data analysis device based on multi-channel video fusion is further configured to: perform video standardization processes including format standardization, filter denoising, video compression, and color space conversion, perform standardization processing on the multi-channel traffic video stream set based on the video standardization processes to obtain a multi-channel standard traffic video stream set, perform traffic flow characteristic analysis on the traffic video monitoring node set to obtain video traffic flow characteristic information, determine a multi-channel video sampling step set according to the video traffic flow characteristic information, and perform downsampling processing on the multi-channel standard traffic video stream set based on the multi-channel video sampling step set to obtain the multi-channel traffic video frame set.

[0053] Further, the video preprocessing module 12 in the traffic data analysis device based on multi-channel video fusion is further configured to: perform initial frame selection on the multi-channel traffic video frame set to obtain an initial multi-channel video frame set, construct an initial multi-channel background model according to the initial multi-channel video frame set, obtain background transformation factors, update the initial multi-channel background model based on the background transformation factors to obtain a multi-channel background updated model, perform foreground object identification on the initial multi-channel video frame set based on the multi-channel background updated model to obtain a multi-channel foreground object model, construct an attention recognition module based on the multi-channel background updated model and the multi-channel foreground object model, perform background recognition and removal on the multi-channel traffic video frame set based on the attention recognition module to obtain the multi-channel traffic foreground image frame set.

[0054] Further, the feature extraction module 13 in the traffic data analysis device based on multi-channel video fusion is further configured to: The traffic feature index set is acquired, and associated feature extraction is performed on the multi-path traffic foreground sequence frame set based on the traffic feature index set, to obtain a multi-path traffic index associated feature set; index calculation is performed on the multi-path traffic index associated feature set according to the traffic feature index set, to determine the multi-path traffic running parameters; abnormal behavior definition is performed on each foreground object in the multi-path foreground object model, to obtain a foreground object abnormal behavior standard rule; data crawling classification and abnormal identification training are performed based on the foreground object abnormal behavior standard rule, to construct a foreground target abnormal identification model; the multi-path traffic foreground sequence frame set is subjected to abnormal target identification based on the foreground target abnormal identification model, to determine the abnormal traffic target information.

[0055] Further, the data fusion module 14 in the traffic data analysis device based on multi-path video fusion is further used for: According to the multi-path traffic video identification code set, multi-path video monitoring space attribute information and multi-path video monitoring time attribute information are determined; a calibration video perspective is selected, the multi-path video monitoring space attribute information is converted and analyzed according to the calibration video perspective, and a multi-path video data transformation matrix set is obtained; the multi-path traffic foreground sequence frame set is mapped and transformed based on the multi-path video data transformation matrix set, to obtain panoramic traffic fusion video information; each image frame in the panoramic traffic fusion video information is subjected to time synchronization alignment based on the multi-path video monitoring time attribute information, to generate the traffic region panoramic video sequence frame.

[0056] Further, the parallel tracking processing module 15 in the traffic data analysis device based on multi-path video fusion is further used for: The multi-path traffic running parameters are subjected to key road node fitting calculation based on the traffic region panoramic video sequence frame, to determine key road node traffic running parameter information; the abnormal traffic target information is subjected to anchor box identification and feature extraction in sequence, to obtain a plurality of abnormal target feature data sets; a feature tracking network is built based on the plurality of abnormal target feature data sets, the plurality of abnormal target feature data sets and the traffic region panoramic video sequence frame are subjected to parallel tracking analysis by using the feature tracking network, to obtain a plurality of abnormal target behavior feature parameters; the target traffic data analysis result is obtained based on the key road node traffic running parameter information and the plurality of abnormal target behavior feature parameters.

[0057] Further, the parallel tracking processing module 15 in the traffic data analysis device based on multi-path video fusion is further used for: Parallel tracking and labeling are performed on the plurality of abnormal target feature data sets and the traffic area panoramic video sequence frames by using the feature tracking network, to obtain a plurality of abnormal target running trajectories; a traffic target abnormal behavior library is constructed based on the foreground object abnormal behavior standard rule; target behavior correlation is performed on the traffic area panoramic video sequence frames according to the plurality of abnormal target running trajectories, to obtain a plurality of abnormal target behavior feature sets; and behavior matching detection is performed on the plurality of abnormal target behavior feature sets based on the traffic target abnormal behavior library, to obtain the plurality of abnormal target behavior feature parameters.

[0058] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The foregoing Figure 1 The traffic data analysis method based on multi-path video fusion in embodiment one and the specific examples are also applicable to the traffic data analysis device based on multi-path video fusion in the present embodiment. Based on the foregoing detailed description of the traffic data analysis method based on multi-path video fusion, those skilled in the art can clearly understand the traffic data analysis device based on multi-path video fusion in the present embodiment. Therefore, for the sake of brevity of the specification, it will not be described in detail here. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant part can be referred to the method part.

[0059] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0060] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application also intends to include these modifications and variations.

Claims

1. A traffic data analysis method based on multi-channel video fusion, characterized in that, include: The target traffic area is obtained, monitoring nodes are deployed in the target traffic area to obtain a set of traffic video monitoring nodes, and a set of multiple traffic video streams is collected through the set of traffic video monitoring nodes. The multi-channel traffic video stream set is subjected to standardization and downsampling processing respectively to obtain a multi-channel traffic video frame set, and the multi-channel traffic video frame set is subjected to background recognition and removal to obtain a multi-channel traffic foreground image frame set. The set of multi-channel traffic foreground image frames is arranged according to time sequence information to obtain a set of multi-channel traffic foreground sequence frames. Traffic data analysis and abnormal target identification are performed on the set of multi-channel traffic foreground sequence frames to determine multi-channel traffic operation parameters and abnormal traffic target information. Spatial encoding is performed on the set of multi-channel traffic foreground sequence frames to obtain a set of multi-channel traffic video identifier codes. The set of multi-channel traffic foreground sequence frames is then aligned and fused according to the set of multi-channel traffic video identifier codes to generate a panoramic video sequence frame of the traffic area. Based on the panoramic video sequence frames of the traffic area, the multi-path traffic operation parameters and abnormal traffic target information are tracked and processed in parallel to obtain the target traffic data analysis results.

2. The traffic data analysis method based on multi-channel video fusion as described in claim 1, characterized in that, The acquisition of the multi-channel traffic video frame set includes: The video standardization process is obtained, which includes format standardization, filtering and noise reduction, video compression, and color space conversion. Based on the video standardization process, the multi-channel traffic video stream set is standardized to obtain a multi-channel standard traffic video stream set. Traffic flow characteristics are analyzed on the set of traffic video monitoring nodes to obtain video traffic flow characteristic information, and a set of multi-channel video sampling step sizes is determined based on the video traffic flow characteristic information. The multi-channel standard traffic video stream set is downsampled based on the multi-channel video sampling step size set to obtain the multi-channel traffic video frame set.

3. The traffic data analysis method based on multi-channel video fusion as described in claim 1, characterized in that, The process of obtaining a set of foreground image frames for multiple traffic routes includes: Initial frames are selected from the set of multiple traffic video frames to obtain an initial set of multiple video frames. Initial background models are then constructed based on the initial set of multiple video frames. Obtain background transformation factors, and update the initial multi-path background model based on the background transformation factors to obtain a multi-path background update model; Based on the multi-path background update model, foreground object identification is performed on the initial multi-path video frame set to obtain a multi-path foreground object model; Based on the multi-path background update model and the multi-path foreground object model, an attention recognition module is constructed. Based on the attention recognition module, background recognition and removal are performed on the multi-path traffic video frame set to obtain the multi-path traffic foreground image frame set.

4. The traffic data analysis method based on multi-channel video fusion as described in claim 3, characterized in that, The determination of multi-route traffic operation parameters and abnormal traffic target information includes: Obtain a set of traffic feature indicators, and extract associated features based on the set of multi-path traffic foreground sequence frames to obtain a set of associated features for multi-path traffic indicators. The multi-route traffic indicator association feature set is calculated according to the traffic characteristic indicator set to determine the multi-route traffic operation parameters; Define abnormal behaviors for each foreground object in the multi-path foreground object model to obtain standard rules for abnormal foreground object behaviors; Based on the aforementioned standard rules for abnormal behavior of foreground objects, data crawling, classification, and abnormal identification training are performed to construct a foreground target abnormal identification model. Based on the aforementioned foreground target anomaly recognition model, abnormal targets are identified in the multi-path traffic foreground sequence frame set to determine the abnormal traffic target information.

5. The traffic data analysis method based on multi-channel video fusion as described in claim 1, characterized in that, The generation of the panoramic video sequence frames of the traffic area includes: Based on the set of multi-channel traffic video identification codes, determine the spatial attribute information and temporal attribute information of the multi-channel video monitoring. Select a calibration video perspective, and perform transformation analysis on the spatial attribute information of the multi-channel video monitoring according to the calibration video perspective to obtain a set of multi-channel video data transformation matrices; Based on the set of transformation matrices for the multi-channel video data, the set of multi-channel traffic foreground sequence frames is mapped, transformed and fused to obtain panoramic traffic fused video information. Based on the time attribute information of the multi-channel video monitoring, the image frames in the panoramic traffic fusion video information are time-synchronized and aligned to generate the panoramic video sequence frame of the traffic area.

6. The traffic data analysis method based on multi-channel video fusion as described in claim 4, characterized in that, The obtained target traffic data analysis results include: Based on the panoramic video sequence frames of the traffic area, the key road node fitting calculation is performed on the multi-road traffic operation parameters to determine the traffic operation parameter information of the key road nodes. The abnormal traffic target information is sequentially subjected to anchor box identification and feature extraction to obtain multiple abnormal target feature datasets; Based on the multiple abnormal target feature datasets, a feature tracking network is built, and the feature tracking network is used to perform parallel tracking analysis on the multiple abnormal target feature datasets and the panoramic video sequence frames of the traffic area to obtain multiple abnormal target behavior feature parameters. Based on the traffic operation parameter information of the key road nodes and the behavioral characteristic parameters of the multiple abnormal targets, the target traffic data analysis results are obtained.

7. The traffic data analysis method based on multi-channel video fusion as described in claim 6, characterized in that, The method of obtaining multiple abnormal target behavior feature parameters includes: The feature tracking network is used to perform parallel tracking and labeling of the multiple abnormal target feature datasets and the panoramic video sequence frames of the traffic area to obtain the running trajectories of multiple abnormal targets; Based on the aforementioned standard rules for abnormal behavior of foreground objects, a database of abnormal behaviors of traffic targets is constructed. Based on the running trajectories of the multiple abnormal targets, the panoramic video sequence frames of the traffic area are correlated with target behaviors to obtain multiple abnormal target behavior feature sets; Based on the traffic target abnormal behavior database, behavior matching detection is performed on the multiple abnormal target behavior feature sets to obtain the multiple abnormal target behavior feature parameters.

8. A traffic data analysis device based on multi-channel video fusion, characterized in that, The step of implementing the traffic data analysis method based on multi-channel video fusion according to any one of claims 1 to 7, wherein the traffic data analysis device based on multi-channel video fusion comprises: The data acquisition module is used to acquire a target traffic area, deploy monitoring nodes in the target traffic area, obtain a set of traffic video monitoring nodes, and acquire a set of multiple traffic video streams through the set of traffic video monitoring nodes. The video preprocessing module is used to perform standardization and downsampling processing on the multi-channel traffic video stream set to obtain a multi-channel traffic video frame set, and to perform background recognition and removal on the multi-channel traffic video frame set to obtain a multi-channel traffic foreground image frame set. The feature extraction module is used to arrange the set of multi-channel traffic foreground image frames according to time sequence information to obtain a set of multi-channel traffic foreground sequence frames, and to perform traffic data analysis and abnormal target identification on the set of multi-channel traffic foreground sequence frames to determine multi-channel traffic operation parameters and abnormal traffic target information. The data fusion module is used to spatially encode the set of multi-channel traffic foreground sequence frames to obtain a set of multi-channel traffic video identifier codes, and to align and fuse the set of multi-channel traffic foreground sequence frames according to the set of multi-channel traffic video identifier codes to generate a panoramic video sequence frame of the traffic area. A parallel tracking processing module is used to perform parallel tracking processing on the multiple traffic operation parameters and abnormal traffic target information based on the panoramic video sequence frames of the traffic area, and obtain the target traffic data analysis results.

Citation Information

Cited By

  • Tunnel traffic flow detection method based on video cloud networking technology

    CN121483046A

  • Intelligent cloud traffic control system and traffic control method

    CN121938204A