Multi-source traffic video data fusion management method and system

By using a multi-source traffic video data fusion management method, continuous spatiotemporal trajectories are generated and intelligent resource allocation is performed, which solves the problems of data silos and inefficient resource allocation, and realizes intelligent traffic management and efficient early warning.

CN120953938BActive Publication Date: 2026-02-06GUANGZHOU TURINGIT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511472986.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-02-06
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Current multi-source traffic video data systems suffer from data silos, making it impossible to form continuous spatiotemporal trajectories, resulting in inconsistent event descriptions, inefficient resource allocation, and difficulty in achieving intelligent management.

Method used

By generating continuous spatiotemporal trajectories of targets through multi-source data association, scoring of viewpoint, occlusion, and sharpness is performed to achieve dynamic resource allocation, and proactive early warning is provided by combining trajectory analysis and behavior recognition.

Benefits of technology

It enables the generation of continuous spatiotemporal trajectories for targets spanning a wide area, improving the intelligence level of traffic management and the efficiency of system operation, and ensuring high-quality analysis and early warning capabilities for critical events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953938B_ABST
    Figure CN120953938B_ABST
Patent Text Reader

Abstract

The application provides a multi-source traffic video data fusion management method and system. The method comprises the following steps: collecting multi-source traffic video data and static metadata, performing target identification, performing associated fusion based on the target identification data and the static metadata, determining associated video sources and generating a target continuous space-time trajectory, performing event detection based on the target identification data and the static metadata, generating standardized semantic event information based on the target continuous space-time trajectory, respectively performing a perspective score on each associated video source based on the standardized semantic event information, respectively performing an occlusion score and an image definition score on each associated video source, and combining the perspective score to perform resource dynamic allocation and output a complete historical trajectory atlas of a target involved in the event, so that collaborative perception, accurate event detection and system resource optimization scheduling of massive heterogeneous traffic monitoring videos can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of traffic monitoring and management, in particular, to a multi-source traffic video data fusion management method and system. BACKGROUND

[0002] With the deepening of the construction of smart cities, the traffic monitoring network is increasingly dense, forming multi-source video data streams from different locations and different types of devices. These video data are complementary in space, time and perspective, and contain rich traffic state and event information. However, the current utilization of multi-source video data has the following technical bottlenecks: 1. Data island problem: each video source usually works independently, lacking an effective correlation and fusion mechanism. When a traffic target (such as a vehicle or a pedestrian) crosses multiple camera fields of view, its trajectory information is fragmented, and it is difficult to form a continuous and complete spatio-temporal trajectory, making it difficult to conduct global behavior analysis; 2. Inconsistent event description: different cameras have different capture angles and clarity for the same traffic event (such as a violation or an accident), resulting in fragmented and low-standardized event description information, making it difficult to generate standardized event reports that can be uniformly understood and processed by machines; 3. Inefficient allocation of system resources: the computing, storage and network bandwidth resources of the monitoring system are limited. Current systems mostly use fixed strategies or simple polling methods to allocate resources, which cannot dynamically and intelligently allocate resources according to the urgency of events and the quality of video sources (such as perspective, obstruction and clarity), resulting in insufficient resources for critical events or waste of resources. Therefore, there is an urgent need in the art for a method that can effectively fuse multi-source heterogeneous video data, automatically generate standardized event information, and intelligently allocate system resources based on perception quality, in order to improve the intelligent level of traffic management and the efficiency of system operation. SUMMARY

[0003] The purpose of the present application is to provide a multi-source traffic video data fusion management method and system, which solves the problem of limited field of view of a single monitoring device through multi-source data correlation, generates continuous spatio-temporal trajectories of targets crossing a wide area, and provides a data basis for macroscopic traffic analysis. At the same time, through multi-dimensional (perspective, obstruction, clarity) perception quality scoring, a dynamic allocation strategy is realized, in which system resources are tilted towards video sources that are better captured, ensuring the quality and efficiency of critical event analysis. And through trajectory analysis and behavior recognition, the behavior of high-risk targets can be actively predicted, and resources can be mobilized in advance for tracking and monitoring, realizing the transition from passive response to active warning.

[0004] The present application also provides a multi-source traffic video data fusion management method, comprising the following steps:

[0005] Collecting multi-source traffic video data and static metadata, and performing target recognition;

[0006] Correlation fusion is performed based on the target recognition data and the static metadata to determine the associated video sources and generate a continuous spatiotemporal trajectory of the target;

[0007] Event detection is performed based on the target recognition data and the static metadata, and standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target;

[0008] A perspective score is respectively assigned to each associated video source based on the standardized semantic event information;

[0009] An occlusion score and an image clarity score are respectively assigned to each associated video source, and resource dynamic allocation is performed in combination with the perspective score to output a complete historical trajectory map of the target involved in the event.

[0010] Optionally, in the multi-source traffic video data fusion management method described in the present application, the multi-source traffic video data and the static metadata are collected, and target recognition is performed, including:

[0011] The multi-source traffic video data is collected, and the static metadata corresponding to each video source is obtained;

[0012] The static metadata includes a device ID, geographic location coordinates, and a timestamp;

[0013] A preset deep learning model is used to perform target recognition on the multi-source traffic video data to obtain target recognition data;

[0014] The target recognition data includes target type data, target attribute data, and corresponding confidence levels.

[0015] Optionally, in the multi-source traffic video data fusion management method described in the present application, the correlation fusion is performed based on the target recognition data and the static metadata to determine the associated video sources and generate a continuous spatiotemporal trajectory of the target, including:

[0016] Based on the target type data and the target attribute data, it is determined whether the traffic targets of different video sources belong to the same physical entity;

[0017] If yes, all video sources corresponding to the same physical entity are marked as the associated video sources of the physical entity;

[0018] The geographic location coordinates in the associated video sources are connected in chronological order to form a continuous spatiotemporal trajectory of the target, which is stored in a historical trajectory database.

[0019] Optionally, in the multi-source traffic video data fusion management method described in the present application, the event detection is performed based on the target recognition data and the static metadata, and the standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target, including:

[0020] Input the target recognition data and static metadata into a preset three-dimensional convolutional neural network model to obtain event type data and generate an event identifier;

[0021] Extract associated target type data and associated target attribute data corresponding to the event from the target recognition data;

[0022] Select the target attribute data with the highest confidence in the associated video source as high-confidence attribute data;

[0023] Encapsulate the event identifier, event type data, associated target type data, high-confidence attribute data, and static metadata corresponding to the event into standardized semantic event information.

[0024] Optionally, in the multi-source traffic video data fusion management method described in the present application, the perspective score of each associated video source based on the standardized semantic event information comprises:

[0025] Process the video frame where the event core region in the standardized semantic event information is located using a preset target detection algorithm to identify the event core region bounding box;

[0026] Calculate the normalized Euclidean distance between the center point of the bounding box and the center point of the video frame, and determine the relative position score according to a preset distance mapping function; calculate the ratio of the area of the bounding box to the total area of the video frame, and determine the proportion score according to a preset area ratio mapping function;

[0027] Sum the relative position score and the proportion score to obtain the perspective score.

[0028] Optionally, in the multi-source traffic video data fusion management method described in the present application, the occlusion score and the image clarity score of each associated video source are calculated, and the resource dynamic allocation and the output of the complete historical trajectory graph of the target involved in the event are combined with the perspective score, comprising:

[0029] Process the video frames of each associated video source through a pre-trained image segmentation model, and count the number of pixel points in the event core region that are classified as occlusion categories;

[0030] Calculate the ratio of the number of occlusion pixel points to the total number of pixel points in the event core region to obtain the occlusion rate and determine the occlusion score;

[0031] Process the event region image block using a preset no-reference image quality assessment algorithm to output an image clarity score;

[0032] Sum the perspective score, occlusion score, and image clarity score to obtain a perception efficiency score;

[0033] According to the perception performance score, the system resources are dynamically allocated, and based on the historical trajectory database, the complete historical trajectory atlas of all targets corresponding to the event identifier is outputted.

[0034] Optionally, in the multi-source traffic video data fusion management method described in the application, further comprising:

[0035] Based on the target continuous spatio-temporal trajectory, the motion parameters of the traffic target are extracted, including the speed sequence, the acceleration sequence and the direction change rate sequence.

[0036] The motion parameters are processed by using a preset behavior recognition model to recognize high-risk driving behaviors.

[0037] If a high-risk driving behavior is recognized, a corresponding behavior label is generated, and the target is marked as a high-risk target.

[0038] Based on the latest standardized semantic event information, the real-time motion state of the high-risk target is obtained, including the real-time geographic location coordinates, the instantaneous speed, the motion direction and the acceleration.

[0039] According to the real-time motion state and the geographic location coordinates of the associated video source, the short-term motion path thereof is predicted.

[0040] The start high-frame-rate mode instruction and the target feature information are sent to the video source on the predicted path.

[0041] In a second aspect, the application provides a multi-source traffic video data fusion management system, which comprises a memory and a processor, wherein the memory stores a multi-source traffic video data fusion management method program, and the multi-source traffic video data fusion management method program is executed by the processor to realize the following steps:

[0042] Multi-source traffic video data and static metadata are collected, and target recognition is performed.

[0043] Based on the target recognition data and the static metadata, associated fusion is performed to determine the associated video sources and generate a target continuous spatio-temporal trajectory.

[0044] Based on the target recognition data and the static metadata, event detection is performed, and standardized semantic event information is generated based on the target continuous spatio-temporal trajectory.

[0045] Based on the standardized semantic event information, a perspective score is respectively given to each associated video source.

[0046] Each associated video source is respectively given an occlusion score and an image clarity score, and the resource dynamic allocation and the complete historical trajectory atlas of the target involved in the event are outputted in combination with the perspective score.

[0047] Optionally, in the multi-source traffic video data fusion management system provided in the application, the multi-source traffic video data and static metadata are collected, and target recognition is performed, comprising:

[0048] The multi-source traffic video data are collected, and static metadata corresponding to each video source are obtained;

[0049] The static metadata comprise device ID, geographic position coordinates and time stamp;

[0050] The preset deep learning model is used to perform target recognition on the multi-source traffic video data, so as to obtain target recognition data;

[0051] The target recognition data comprise target type data, target attribute data and corresponding confidence.

[0052] Optionally, in the multi-source traffic video data fusion management system provided in the application, the target recognition data and the static metadata are associated and fused to determine associated video sources and generate a target continuous space-time trajectory, comprising:

[0053] Based on the target type data and the target attribute data, it is determined whether the traffic targets of different video sources belong to the same physical entity;

[0054] If yes, all video sources corresponding to the same physical entity are marked as the associated video sources of the physical entity;

[0055] The geographic position coordinates in the associated video sources are connected in time stamp order to form a target continuous space-time trajectory, and the target continuous space-time trajectory is stored in a historical trajectory database.

[0056] As can be seen from the above, the multi-source traffic video data fusion management method and system provided in the application solve the problem of limited field of view of a single monitoring device through multi-source data association, generate a continuous space-time trajectory of a target across a wide area, and provide a data basis for macro traffic analysis. Meanwhile, through multi-dimensional (perspective, occlusion, definition) perception quality scoring, a dynamic allocation strategy is realized, in which system resources are inclined to video sources that can capture better, so as to guarantee the quality and efficiency of key event analysis. Through trajectory analysis and behavior recognition, the behavior of a high-risk target can be actively predicted, and resources can be mobilized in advance for tracking and monitoring, so that a change from passive response to active early warning is realized.

[0057] Other features and advantages of the application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the application as described in the written description and claims. The purposes and other advantages of the application will be realized and attained by the structures particularly pointed out in the written description and claims. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments of the present application will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0059] Figure 1 The flow chart of the multi-source traffic video data fusion management method provided in the embodiments of the present application;

[0060] Figure 2 The flow chart of the multi-source traffic video data fusion management method provided in the embodiments of the present application for generating standardized semantic event information;

[0061] Figure 3 The flow chart of the multi-source traffic video data fusion management method provided in the embodiments of the present application for performing view angle scoring;

[0062] Figure 4 The flow chart of the multi-source traffic video data fusion management method provided in the embodiments of the present application for performing resource dynamic allocation. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely in the embodiments of the present application in combination with the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present application.

[0064] It should be noted that similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0065] Please refer to Figure 1 , Figure 1 The flow chart of the multi-source traffic video data fusion management method in some embodiments of the present application. The multi-source traffic video data fusion management method is used in a terminal device, such as a computer, a mobile phone terminal, etc. The multi-source traffic video data fusion management method includes the following steps:

[0066] S11, collect multi-source traffic video data and static metadata, and perform target recognition;

[0067] S12, perform correlation fusion based on the target recognition data and the static metadata, determine the correlated video sources, and generate a continuous spatiotemporal trajectory of the target;

[0068] S13, perform event detection based on the target recognition data and the static metadata, and generate standardized semantic event information based on the continuous spatiotemporal trajectory of the target;

[0069] S14, score the perspective of each correlated video source based on the standardized semantic event information;

[0070] S15, score the occlusion and the image clarity of each correlated video source, and combine the perspective score to perform dynamic allocation of resources and output a complete historical trajectory map of the target involved in the event.

[0071] It should be noted that, first, multi-source traffic video data and static metadata are collected, and a deep learning model is used for target recognition to extract target and attribute information. Second, based on the recognition results and metadata, data correlation fusion is performed to determine whether the targets in different video sources are the same physical entity, thereby determining the correlated video sources and generating a continuous spatiotemporal trajectory of the target across multiple cameras. Next, event detection is performed, and using the generated continuous trajectory and other information, a standardized semantic event description is constructed containing high-value information such as event identification, type, and target attributes. Then, based on the standardized event information, the perspective quality of each correlated video source capturing the event is evaluated (perspective score). Finally, the occlusion and image clarity of each video source are further evaluated (occlusion score and clarity score), and the comprehensive quality is evaluated in combination with the perspective score, and the system resources (such as computing resources, storage priority, bandwidth) are dynamically allocated accordingly, and a complete historical trajectory map of the target involved in the event is output. Through the above method, the invention realizes the integration of data fusion, semantic description, quality evaluation, and resource scheduling, effectively solves the problems of data silos, information fragmentation, and inefficient resource allocation, and significantly improves the perception ability and operation efficiency of the traffic monitoring system.

[0072] According to the embodiment of the present application, the multi-source traffic video data and static metadata are collected, and the target recognition is performed, comprising:

[0073] Collecting multi-source traffic video data and obtaining static metadata corresponding to each video source;

[0074] The static metadata includes device ID, geographic location coordinates, and timestamp;

[0075] A preset deep learning model is used to perform target recognition on the multi-source traffic video data to obtain target recognition data;

[0076] The target recognition data includes target type data, target attribute data and corresponding confidence thereof.

[0077] It should be noted that the target recognition data is obtained by using a preset deep learning model (such as a single detection algorithm (YOLO, You Only Look Once), a single multi-frame detection algorithm (SSD, Single Shot MultiBox Detector), and a fast region-based convolutional neural network (Faster R-CNN, Faster Region-based Convolutional Neural Network)) to recognize targets in multi-source traffic video data. The target attribute data is used to finely describe the recognized targets, and the specific content thereof depends on the target type. For motor vehicles, the target attribute data includes but is not limited to license plate number, vehicle color, vehicle brand, vehicle model and vehicle type; for pedestrians and non-motor vehicles, the target attribute data includes but is not limited to subject color, gender, approximate age group and accessory information carried. In addition, general attributes such as the motion speed, motion direction and position information of the target in the image are also included. Each attribute data corresponds to a confidence generated by the recognition model, which is used to measure the reliability of the attribute recognition.

[0078] According to the embodiment of the present application, the target recognition data and the static metadata are associated and fused to determine the associated video source and generate the continuous spatio-temporal trajectory of the target, which comprises:

[0079] Based on the target type data and the target attribute data, it is determined whether the traffic targets of different video sources belong to the same physical entity;

[0080] If yes, all video sources corresponding to the same physical entity are marked as the associated video source of the physical entity;

[0081] The geographical position coordinates in the associated video source are connected in chronological order to form the continuous spatio-temporal trajectory of the target, and stored in the historical trajectory database.

[0082] It should be noted that the target type data and target attribute data are used to determine whether the traffic targets of different video sources belong to the same physical entity, which is achieved through a multi-level correlation matching process, including: 1. Space mapping and preliminary screening: According to the geographic location coordinates in the static metadata of each video source, all target recognition data is mapped to a unified electronic map coordinate system to form a global view based on geographic space, and a set of targets that may be associated in space is preliminarily screened out. 2. Spatiotemporal continuity constraint analysis: According to the timestamp of the target and its position change based on the target in the continuous video frames and the timestamp information, the motion parameters including instantaneous speed, motion direction and acceleration are calculated to construct a spatiotemporal continuity constraint model, and the coherence of the targets in different events in time sequence and motion trajectory is calculated, and based on this, the association probability of them being the same physical entity is calculated, and the candidate targets that are not coherent in space and time are excluded. 3. Multi-feature fusion precise matching: The appearance features (such as color, vehicle type) and identity features (such as license plate number) of the traffic targets are integrated, and a preset multi-feature similarity calculation model is used for weighted fusion and association judgment. Finally, based on the fusion similarity score and the spatiotemporal association probability, it is determined whether the targets in different video sources belong to the same physical entity.

[0083] Please refer to Figure 2 , Figure 2 is a flowchart of a method for generating standardized semantic event information in some embodiments of the multi-source traffic video data fusion management method. According to the embodiments of the present application, event detection is performed based on target recognition data and static metadata, and standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target, including:

[0084] S21, inputting the target recognition data and static metadata into a preset three-dimensional convolutional neural network model to obtain event type data, and generating an event identifier;

[0085] S22, extracting the associated target type data and its associated target attribute data corresponding to the event from the target recognition data;

[0086] S23, selecting the target attribute data with the highest confidence in the associated video source as high-confidence attribute data;

[0087] S24, encapsulating the event identifier, event type data, associated target type data, high-confidence attribute data, and event corresponding static metadata as standardized semantic event information.

[0088] It should be noted that in the event detection step, the preset three-dimensional convolutional neural network model can analyze the spatiotemporal features of the video stream, thereby perceiving traffic events (traffic accidents, traffic violations, high-risk driving behaviors, etc.) and performing event encapsulation.

[0089] Please refer toFigure 3 , Figure 3 is a flowchart of the view angle scoring of the multi-source traffic video data fusion management method in some embodiments of the present application. According to the embodiment of the present application, the view angle scoring of each associated video source based on the standardized semantic event information comprises:

[0090] S31, processing the video frame in which the event core region is located in the standardized semantic event information by using a preset target detection algorithm, and identifying the event core region bounding box;

[0091] S32, calculating the normalized Euclidean distance between the center point of the bounding box and the center point of the video frame, and determining the relative position score according to a preset distance mapping function; S33, calculating the ratio of the area of the bounding box to the total area of the video frame, and determining the proportion score according to a preset area proportion mapping function;

[0092] S34, weighting and summing the relative position score and the proportion score to obtain the view angle score.

[0093] It should be noted that the preset target detection algorithm (YOLO model) is used to process the video frame where the event core region in the standardized semantic event information is located, to generate an initial candidate bounding box; at the same time, the traffic event type data, the associated target type data, and the high-confidence attribute data (target position coordinates, size, etc.) are used as constraint conditions to screen and calibrate the initial candidate bounding box (such as filtering the bounding box without target type based on the associated target type, excluding irrelevant targets in combination with the event type characteristics, adjusting the center position of the bounding box based on the target coordinates, and correcting the range of the bounding box according to the target size), to finally obtain the event core region bounding box. For example, taking a "motor vehicle rear-end collision accident" as an example, the standardized semantic event information thereof includes a traffic event type (motor vehicle rear-end collision accident, requiring at least 2 motor vehicles in close contact), an associated target type (small car A and small car B, both of which are "motor vehicle" category), high-confidence attribute data (car A position coordinates (300, 200), size 150x80 pixels, car B position coordinates (350, 200), size 140x75 pixels, confidence greater than or equal to 0.9), and a video frame where the event core region is located (key frame at the accident time positioned by timestamp); the preset target detection algorithm generates an initial candidate bounding box after processing the video frame, including the bounding boxes corresponding to small car A, small car B, roadside pedestrians, distant trucks, and road barriers; then, the bounding boxes containing "pedestrian" and "static facility" and the truck bounding box of non-"small car" type are filtered out as constraints of the associated target type, and only the bounding boxes of car A and B are reserved; then, the two-car bounding boxes are confirmed to meet the condition of "≥2 motor vehicles and distance ≤50 pixels (overlap exists, which meets the close contact feature)" in combination with the traffic event type feature screening; finally, the high-confidence attribute data is calibrated based on the center (325, 200) of the core region determined by the centers (300, 200) and (350, 200) of the two cars, and the bounding box is expanded according to the size of the two cars and the collision impact range (width covers 205-460 pixels, height covers 150-250 pixels), to finally obtain the event core region bounding box containing the two cars and the collision impact area, which accurately matches the constraint requirements of the event type, the associated target type, and the high-confidence attribute data.

[0094] The distance mapping function is a mathematical rule or calculation rule for converting the "distance of the event area from the center of the picture" into a "relative position score". Specifically, a pre-set mapping rule is applied to the calculated normalized Euclidean distance D (the value is between 0 and 1, 0 represents the center, and 1 represents the edge of the picture), and the rule is usually "the smaller the distance, the higher the score", and a relative position score between 0 and 1 is output (for example, 0.95 points). For example, define a simple linear mapping function: score = 1 - distance. If the event is exactly in the center (distance = 0), the score = 1 - 0 = 1.0 (full score); if the event is at the edge of the picture (distance = 0.8), the score = 1 - 0.8 = 0.2 (low score); if the event is in the middle position (distance = 0.3), the score = 1 - 0.3 = 0.7 (medium score).

[0095] The area ratio mapping function is a mathematical rule for converting the "area ratio of the event area in the picture" into a "proportion score". Specifically, a pre-set mapping rule is applied to the ratio R (the value is between 0 and 1) of the calculated core area of the event to the total area of the video frame, and a proportion score between 0 and 1 is output. The rule is usually a "moderate" preference, that is, the proportion cannot be too small (details cannot be seen) or too large (the global context is lost). For example, define a "bell-shaped" mapping function that prefers a proportion between 30% and 60%. If the event area proportion is too small (R = 0.1), the score is very low (0.2) because it cannot be seen; if the event area proportion is moderate (R = 0.4), the score is very high (0.95); if the event area proportion is too large (R = 0.9), the score will also decrease (0.4) because only the local part is captured and the surrounding environment cannot be seen.

[0096] Please refer to Figure 4 , Figure 4 is a flowchart of resource dynamic allocation of a multi-source traffic video data fusion management method in some embodiments of the present application. According to the embodiment of the present application, the occlusion score and the image clarity score of each associated video source are calculated, and the resource dynamic allocation is performed in combination with the view angle score to output the complete historical trajectory graph of the target involved in the event, including:

[0097] S41, processing the video frames of each associated video source by a pre-trained image segmentation model, and counting the number of pixel points in the event core area classified as the occlusion category;

[0098] S42, calculating the ratio of the number of occlusion pixel points to the total number of pixel points in the event core area to obtain the occlusion rate, and determining the occlusion score;

[0099] S43, processing the event region image block by using a preset no-reference image quality evaluation algorithm to output an image definition score;

[0100] S44, weighting and summing the view angle score, the occlusion score and the image definition score to obtain a perception performance score;

[0101] S45, dynamically allocating system resources according to the perception performance score and outputting complete historical trajectory atlas of all targets corresponding to the event identifier based on the historical trajectory database and taking the event identifier as an index.

[0102] It should be noted that the determination method of the occlusion score is as follows: the system quantifies the severity of occlusion by calculating the obtained occlusion rate, in the application embodiment, a nonlinear scoring mapping function is adopted, the lower the occlusion rate, the higher the score, and higher weight sensitivity is given to low occlusion. Specifically, when the occlusion rate is close to 0 (no occlusion), the occlusion score should be close to full score (for example, 1.0), to reflect its excellent observation quality; as the occlusion rate increases, the score should accelerate to decline; when the occlusion rate exceeds a certain threshold (for example, 70%), the score should tend to 0, indicating that the video source has almost lost the observation value due to serious occlusion.

[0103] According to the perception performance score, the system resources are dynamically allocated, and the computing, network bandwidth and storage resources of the dynamic allocation system are dynamically allocated. Specifically, an instruction is issued to the front-end device or transmission node of the video source with a high perception performance score, requesting or receiving a high-definition, lossless video stream; an instruction is issued to the front-end device or transmission node of the video source with a low perception performance score, to reduce the transmission code rate or resolution of the video stream. The video data generated by the video source with a high perception performance score triggers an event-associated storage strategy, and is stored with a high code rate and a long period, and is additionally labeled with event metadata; the video data generated by the video source with a low perception performance score is stored by using a conventional short-period rolling coverage storage strategy.

[0104] According to the embodiment of the application, the method further comprises:

[0105] Based on the target continuous spatio-temporal trajectory, the motion parameters of the traffic target are extracted, including a speed sequence, an acceleration sequence and a direction change rate sequence;

[0106] The motion parameters are processed by using a preset behavior recognition model to recognize high-risk driving behaviors;

[0107] If a high-risk driving behavior is recognized, a corresponding behavior label is generated, and the target is marked as a high-risk target;

[0108] Based on the latest standardized semantic event information, the real-time motion state of the high-risk target is obtained, including a real-time geographic position coordinate, an instantaneous speed, a motion direction and an acceleration.

[0109] predicting a short-term motion path according to the real-time motion state and the geographical position coordinates of the associated video source;

[0110] sending a start high frame rate mode instruction and target feature information to the video source on the predicted path.

[0111] It should be noted that by predicting the short-term motion path and mobilizing resources (starting the high frame rate mode) in advance, the system can be arranged in advance to ensure that the high-risk target can also be continuously captured clearly on the subsequent path, avoiding the problems of losing or losing key details, and providing a complete evidence chain for subsequent disposal.

[0112] The application also discloses a multi-source traffic video data fusion management system, comprising a memory and a processor, wherein the memory stores a multi-source traffic video data fusion management method program, and the multi-source traffic video data fusion management method program is executed by the processor to realize the following steps:

[0113] collecting multi-source traffic video data and static metadata, and performing target identification;

[0114] correlation fusion based on the target identification data and the static metadata, determining associated video sources and generating a target continuous space-time trajectory;

[0115] event detection based on the target identification data and the static metadata, and generating standardized semantic event information based on the target continuous space-time trajectory;

[0116] scoring the view angle of each associated video source based on the standardized semantic event information;

[0117] scoring the occlusion and image clarity of each associated video source, and combining the view angle score to perform dynamic resource allocation and output a complete historical trajectory map of the target involved in the event.

[0118] It should be noted that first, multi-source traffic video data and static metadata are collected, and a deep learning model is used for target recognition to extract target and attribute information. Second, based on the recognition results and metadata, data correlation fusion is performed to determine whether the targets in different video sources are the same physical entity, thereby determining the associated video sources and generating continuous spatio-temporal trajectories of the targets across multiple cameras. Next, event detection is performed, and the generated continuous trajectories and other information are used to construct standardized semantic event descriptions containing high-value information such as event identification, type, and target attributes. Then, based on the standardized event information, the perspective quality of each associated video source capturing the event is evaluated (perspective score). Finally, the occlusion and image clarity of each video source are further evaluated (occlusion score and clarity score), and the perspective score is combined for comprehensive quality evaluation, based on which system resources (such as computing resources, storage priority, and bandwidth) are dynamically allocated, and the complete historical trajectory map of the targets involved in the event is output. Through the above method, the invention realizes the integration of data fusion, semantic description, quality evaluation, and resource scheduling, effectively solving the problems of data silos, information fragmentation, and inefficient resource allocation, and significantly improving the perception ability and operation efficiency of the traffic monitoring system.

[0119] According to the embodiment of the present application, the multi-source traffic video data and static metadata are collected and target recognition is performed, comprising:

[0120] Collecting multi-source traffic video data and obtaining static metadata corresponding to each video source;

[0121] The static metadata includes device ID, geographic location coordinates, and timestamp.

[0122] Using a preset deep learning model to perform target recognition on the multi-source traffic video data to obtain target recognition data;

[0123] The target recognition data includes target type data, target attribute data, and their corresponding confidence.

[0124] It should be noted that the preset deep learning model (such as single detection algorithm (YOLO, You Only Look Once), single multi-frame detection algorithm (SSD, Single Shot MultiBox Detector), and fast region-based convolutional neural network (Faster R-CNN, Faster Region-based Convolutional Neural Network)) is used to identify the target in the multi-source traffic video data, and target identification data is obtained. The target attribute data is used to finely describe the identified target, and the specific content depends on the target type. For motor vehicles, the target attribute data includes but is not limited to license plate number, vehicle color, vehicle brand, vehicle model and vehicle type; for pedestrians and non-motor vehicles, the target attribute data includes but is not limited to subject color, gender, approximate age group and accessory information carried. In addition, it also includes the general attributes such as the motion speed, the motion direction and the position information in the image of the target. Each attribute data corresponds to a confidence degree generated by the identification model, which is used to measure the reliability of the attribute identification.

[0125] According to the embodiment of the present application, the association fusion based on the target identification data and the static metadata is performed to determine the associated video source and generate the continuous spatio-temporal trajectory of the target, which comprises:

[0126] Based on the target type data and the target attribute data, it is judged whether the traffic target of different video sources belongs to the same physical entity;

[0127] If yes, all video sources corresponding to the same physical entity are marked as the associated video source of the physical entity;

[0128] The geographical position coordinates in the associated video source are connected in the time stamp order to form the continuous spatio-temporal trajectory of the target, and are stored in the historical trajectory database.

[0129] It should be noted that the target type data and the target attribute data are used to determine whether the traffic targets of different video sources belong to the same physical entity, which is specifically implemented through a multi-level correlation matching process, including: 1. Space mapping and preliminary screening: According to the geographic position coordinates in the static metadata of each video source, all target recognition data is mapped to a unified electronic map coordinate system to form a global view based on geographic space, and a target set that may be correlated in space is preliminarily screened out. 2. Spatiotemporal continuity constraint analysis: According to the time stamp of the target and the position change of the target in the continuous video frame and the time stamp information, the motion parameters including instantaneous speed, motion direction and acceleration are calculated to construct a spatiotemporal continuity constraint model, the coherence of the targets in different events in time sequence and motion trajectory is calculated, and the correlation probability of the targets being the same physical entity is calculated based on this, and the candidate targets that are not coherent in space and time are excluded. 3. Multi-feature fusion accurate matching: The appearance features (such as color, vehicle type) and identity features (such as license plate number) of the traffic targets are comprehensively fused and correlated by a preset multi-feature similarity calculation model. Finally, based on the fusion similarity score and the spatiotemporal correlation probability, it is determined whether the targets in different video sources belong to the same physical entity.

[0130] According to the embodiment of the application, the event detection is performed based on the target recognition data and the static metadata, and the standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target, including:

[0131] The target recognition data and the static metadata are input into a preset three-dimensional convolutional neural network model to obtain event type data, and an event identifier is generated;

[0132] The associated target type data and the associated target attribute data corresponding to the event are extracted from the target recognition data;

[0133] The target attribute data with the highest confidence in the associated video source is selected as high-confidence attribute data;

[0134] The event identifier, the event type data, the associated target type data, the high-confidence attribute data, and the static metadata corresponding to the event are collectively encapsulated as standardized semantic event information.

[0135] It should be noted that in the event detection step, the preset three-dimensional convolutional neural network model can analyze the spatiotemporal features of the video stream, thereby perceiving traffic events (traffic accidents, traffic violations, high-risk driving behaviors, etc.) and performing event encapsulation.

[0136] According to the embodiment of the application, the standardized semantic event information is used to score the perspective of each associated video source, including:

[0137] The preset target detection algorithm is used to process a video frame where an event core region in the standardized semantic event information is located, and a bounding box of the event core region is recognized;

[0138] A normalized Euclidean distance between the center point of the bounding box and the center point of the video frame is calculated, and a relative position score is determined according to a preset distance mapping function;

[0139] A ratio of the area of the bounding box to the total area of the video frame is calculated, and a proportion score is determined according to a preset area proportion mapping function;

[0140] The relative position score and the proportion score are weighted and summed to obtain a view angle score.

[0141] It should be noted that the preset target detection algorithm (YOLO model) is used to process the video frame where the event core region in the standardized semantic event information is located, to generate an initial candidate bounding box; at the same time, the traffic event type data, the associated target type data, and the high-confidence attribute data (target position coordinates, size, etc.) are used as constraint conditions to screen and calibrate the initial candidate bounding box (such as filtering the bounding box without target type based on the associated target type, excluding irrelevant targets in combination with the event type characteristics, adjusting the center position of the bounding box based on the target coordinates, and correcting the range of the bounding box according to the target size), and finally obtaining the event core region bounding box. For example, taking a "motor vehicle rear-end collision accident" as an example, the standardized semantic event information thereof includes a traffic event type (motor vehicle rear-end collision accident, requiring at least 2 motor vehicles in close contact), an associated target type (small car A and small car B, both of which are "motor vehicle" category), high-confidence attribute data (car A position coordinates (300, 200), size 150x80 pixels, car B position coordinates (350, 200), size 140x75 pixels, confidence greater than or equal to 0.9), and a video frame where the event core region is located (key frame at the accident time positioned by timestamp); the preset target detection algorithm generates an initial candidate bounding box after processing the video frame, including the bounding boxes corresponding to small car A, small car B, roadside pedestrians, distant trucks, and road barriers; then, the bounding boxes containing "pedestrian" and "static facility" and the truck bounding box of non-"small car" type are filtered out as constraints of the associated target type, and only the bounding boxes of car A and B are reserved; then, the two-car bounding boxes are confirmed to meet the condition of "≥2 motor vehicles and distance ≤50 pixels (overlap exists, which meets the close contact feature)" in combination with the traffic event type feature screening; finally, the high-confidence attribute data is calibrated based on the center (325, 200) of the core region determined by the centers (300, 200) and (350, 200) of the two cars, and the bounding box is expanded according to the size of the two cars and the collision impact range (width covers 205-460 pixels, height covers 150-250 pixels), and finally the event core region bounding box containing the two cars and the collision impact area is obtained, which accurately matches the constraint requirements of the event type, the associated target type, and the high-confidence attribute data.

[0142] The distance mapping function is a mathematical rule or calculation rule for converting the "distance of the event area from the center of the picture" into a "relative position score". Specifically, a pre-set mapping rule is applied to the calculated normalized Euclidean distance D (the value is between 0 and 1, 0 represents the center, and 1 represents the edge of the picture), which is usually "the smaller the distance, the higher the score", and a relative position score between 0 and 1 is output (for example, 0.95 points). For example, define a simple linear mapping function: score = 1 - distance. If the event is exactly in the center (distance = 0), the score = 1 - 0 = 1.0 (full score); if the event is at the edge of the picture (distance = 0.8), the score = 1 - 0.8 = 0.2 (low score); if the event is in the middle (distance = 0.3), the score = 1 - 0.3 = 0.7 (medium score).

[0143] The area ratio mapping function is a mathematical rule for converting the "area ratio of the event area in the picture" into a "proportion score". Specifically, a pre-set mapping rule is applied to the ratio R (the value is between 0 and 1) of the calculated core area of the event area to the total area of the video frame, and a proportion score between 0 and 1 is output. This rule usually has a "moderate" preference, that is, the proportion cannot be too small (details cannot be seen) or too large (the global context is lost). For example, define a "bell-shaped" mapping function that prefers a proportion between 30% and 60%. If the event area proportion is too small (R = 0.1), the score is very low (0.2) because it cannot be seen; if the event area proportion is moderate (R = 0.4), the score is very high (0.95); if the event area proportion is too large (R = 0.9), the score will also decrease (0.4) because only the local part is captured and the surrounding environment cannot be seen.

[0144] According to an embodiment of the present application, the respective associated video sources are subjected to occlusion scoring and image clarity scoring, and the view angle score is combined to perform resource dynamic allocation and output a complete historical trajectory map of the target involved in the event, comprising:

[0145] The video frames of each associated video source are processed by a pre-trained image segmentation model, and the number of pixel points in the event core area classified as an occlusion category is counted;

[0146] The ratio of the number of occlusion pixel points to the total number of pixel points in the event core area is calculated to obtain an occlusion rate, and an occlusion score is determined;

[0147] The event area image block is processed using a pre-set no-reference image quality assessment algorithm to output an image clarity score;

[0148] The view angle score, the occlusion score and the image clarity score are weighted and summed to obtain a perception performance score;

[0149] According to the perception efficiency score, system resources are dynamically allocated, and based on the historical trajectory database, all complete historical trajectory graphs of the target corresponding to the event identifier are outputted.

[0150] It should be noted that the determination method of the occlusion score is as follows: the system quantifies the severity of occlusion by calculating the occlusion rate, in the application embodiment, a nonlinear scoring mapping function is adopted, the lower the occlusion rate, the higher the score, and higher weight sensitivity is given to low occlusion. Specifically, when the occlusion rate is close to 0 (no occlusion), the occlusion score should be close to full score (for example, 1.0), to reflect its excellent observation quality; as the occlusion rate increases, the score should accelerate to decline; when the occlusion rate exceeds a certain threshold (for example, 70%), the score should tend to 0, indicating that the video source has almost lost its observation value due to severe occlusion.

[0151] According to the perception efficiency score, system resources are dynamically allocated, and the computing, network bandwidth and storage resources of the dynamic allocation system. Specifically, instructions are issued to the front-end device or transmission node of the video source with high perception efficiency score, requesting or receiving high-definition, lossless video streams; instructions are issued to the front-end device or transmission node of the video source with low perception efficiency score, reducing the transmission code rate or resolution of the video stream. For the video data generated by the video source with high perception efficiency score, trigger the event association storage strategy to store with high code rate and long period, and add event metadata label; for the video data generated by the video source with low perception efficiency score, adopt the conventional short-period rolling coverage storage strategy.

[0152] According to the embodiment of the application, further comprising:

[0153] Based on the target continuous spatio-temporal trajectory, the motion parameters of the traffic target are extracted, including the speed sequence, the acceleration sequence and the direction change rate sequence;

[0154] The motion parameters are processed by using a preset behavior recognition model to recognize high-risk driving behaviors;

[0155] If a high-risk driving behavior is recognized, a corresponding behavior label is generated, and the target is marked as a high-risk target;

[0156] Based on the latest standardized semantic event information, the real-time motion state of the high-risk target is obtained, including the real-time geographic location coordinates, the instantaneous speed, the motion direction and the acceleration;

[0157] According to the real-time motion state and the geographic location coordinates of the associated video source, the short-term motion path thereof is predicted;

[0158] The start high-frame-rate mode instruction and the target feature information are sent to the video source on the predicted path.

[0159] It should be noted that by predicting the short-term motion path and mobilizing resources in advance (starting the high frame rate mode), the system can be laid out in advance to ensure that high-risk targets can also be continuously captured clearly on the subsequent path, avoiding the problem of losing or missing key details, and providing a complete evidence chain for subsequent disposal.

[0160] The multi-source traffic video data fusion management method and system disclosed in the application solves the problem of limited field of view of a single monitoring device through multi-source data association, generates a continuous space-time trajectory of a target across a wide area, and provides a data basis for macro traffic analysis. At the same time, through multi-dimensional (perspective, occlusion, definition) perception quality scoring, a dynamic allocation strategy is realized, in which system resources are tilted towards video sources that are better captured, ensuring the quality and efficiency of key event analysis. And through trajectory analysis and behavior recognition, the behavior of high-risk targets can be actively predicted, and resources can be mobilized in advance for tracking and monitoring, realizing the transition from passive response to active warning.

[0161] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0162] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units; they can be located in one place or distributed on multiple network units; part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0163] In addition, each functional unit in each embodiment of the application can be integrated into one processing unit, or each unit can be a unit alone, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.

Claims

1. A multi-source traffic video data fusion management method, characterized in that, Includes the following steps: Collect multi-source traffic video data and static metadata, and perform target recognition; Based on the correlation and fusion of target recognition data and static metadata, the associated video source is determined and the continuous spatiotemporal trajectory of the target is generated. Event detection is performed based on target recognition data and static metadata, and standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target. The video frames containing the core region of the event in the standardized semantic event information are processed using a preset target detection algorithm to identify the bounding box of the core region of the event. Calculate the normalized Euclidean distance between the center point of the bounding box and the center point of the video frame, and determine the relative position score according to the preset distance mapping function; Calculate the ratio of the bounding box area to the total area of ​​the video frame, and determine the percentage score based on a preset area ratio mapping function; The relative position score and the proportion score are weighted and summed to obtain the perspective score; The video frames of each associated video source are processed by a pre-trained image segmentation model, and the number of pixels in the core area of ​​the event that are classified as occlusions is counted. Calculate the ratio of the number of pixels of the occluded object to the total number of pixels in the core area of ​​the event to obtain the occlusion rate and determine the occlusion score; The image patch of the event region is processed using a pre-defined no-reference image quality assessment algorithm, and an image sharpness score is output. The perceptual performance score is obtained by weighted summation of the viewpoint score, occlusion score, and image sharpness score. Based on the perception performance score, the system resources are dynamically allocated, and based on the historical trajectory database, using the event identifier as an index, the complete historical trajectory map of all targets corresponding to the event identifier is output.

2. The multi-source traffic video data fusion management method according to claim 1, characterized in that, The process of collecting multi-source traffic video data and static metadata, and performing target recognition, includes: Collect multi-source traffic video data and obtain static metadata corresponding to each video source; The static metadata includes device ID, geographic location coordinates, and timestamp; A pre-defined deep learning model is used to perform target recognition on multi-source traffic video data to obtain target recognition data. Target identification data includes target type data, target attribute data, and their corresponding confidence levels.

3. The multi-source traffic video data fusion management method according to claim 2, characterized in that, The process of associating and fusing target recognition data and static metadata to determine associated video sources and generate continuous spatiotemporal trajectories of the target includes: Based on target type data and target attribute data, determine whether traffic targets from different video sources belong to the same physical entity; If so, all video sources corresponding to the same physical entity will be marked as associated video sources of that physical entity; The geographic coordinates in the associated video sources are connected in time stamp order to form a continuous spatiotemporal trajectory of the target, which is then stored in the historical trajectory database.

4. The multi-source traffic video data fusion management method according to claim 3, characterized in that, The event detection based on target recognition data and static metadata, and the generation of standardized semantic event information based on the continuous spatiotemporal trajectory of the target, includes: The target identification data and static metadata are input into a preset three-dimensional convolutional neural network model for processing to obtain event type data and generate event identifiers. Extract the associated target type data and associated target attribute data corresponding to the event from the target identification data; Select the target attribute data with the highest confidence among the associated video sources as high-confidence attribute data; The event identifier, event type data, associated target type data, high confidence attribute data, and the corresponding static metadata of the event are collectively encapsulated into standardized semantic event information.

5. The multi-source traffic video data fusion management method according to claim 4, characterized in that, Also includes: Based on the continuous spatiotemporal trajectory of the target, the motion parameters of the traffic target are extracted, including velocity sequence, acceleration sequence and direction change rate sequence; The motion parameters are processed using a preset behavior recognition model to identify high-risk driving behaviors; If high-risk driving behavior is identified, a corresponding behavior label is generated, and the target is marked as a high-risk target. Real-time motion status of high-risk targets is obtained based on the latest standardized semantic event information, including real-time geographic location coordinates, instantaneous velocity, direction of motion, and acceleration; Predict its short-term motion path based on its real-time motion status and the geographical coordinates of its associated video sources; Send a command to activate high frame rate mode and target feature information to the video source on the prediction path.

6. A multi-source traffic video data fusion management system, characterized in that, The system includes a memory and a processor. The memory stores a program for a multi-source traffic video data fusion management method. When the program for the multi-source traffic video data fusion management method is executed by the processor, it performs the following steps: Collect multi-source traffic video data and static metadata, and perform target recognition; Based on the correlation and fusion of target recognition data and static metadata, the associated video source is determined and the continuous spatiotemporal trajectory of the target is generated. Event detection is performed based on target recognition data and static metadata, and standardized semantic event information is generated based on the continuous spatiotemporal trajectory of the target. The video frames containing the core region of the event in the standardized semantic event information are processed using a preset target detection algorithm to identify the bounding box of the core region of the event. Calculate the normalized Euclidean distance between the center point of the bounding box and the center point of the video frame, and determine the relative position score according to the preset distance mapping function; Calculate the ratio of the bounding box area to the total area of ​​the video frame, and determine the percentage score based on a preset area ratio mapping function; The relative position score and the proportion score are weighted and summed to obtain the perspective score; The video frames of each associated video source are processed by a pre-trained image segmentation model, and the number of pixels in the core area of ​​the event that are classified as occlusions is counted. Calculate the ratio of the number of pixels of the occluded object to the total number of pixels in the core area of ​​the event to obtain the occlusion rate and determine the occlusion score; The image patch of the event region is processed using a pre-defined no-reference image quality assessment algorithm, and an image sharpness score is output. The perceptual performance score is obtained by weighted summation of the viewpoint score, occlusion score, and image sharpness score. Based on the perception performance score, the system resources are dynamically allocated, and based on the historical trajectory database, using the event identifier as an index, the complete historical trajectory map of all targets corresponding to the event identifier is output.

7. The multi-source traffic video data fusion management system according to claim 6, characterized in that, The process of collecting multi-source traffic video data and static metadata, and performing target recognition, includes: Collect multi-source traffic video data and obtain static metadata corresponding to each video source; The static metadata includes device ID, geographic location coordinates, and timestamp; A pre-defined deep learning model is used to perform target recognition on multi-source traffic video data to obtain target recognition data. Target identification data includes target type data, target attribute data, and their corresponding confidence levels.

8. The multi-source traffic video data fusion management system according to claim 7, characterized in that, The process of associating and fusing target recognition data and static metadata to determine associated video sources and generate continuous spatiotemporal trajectories of the target includes: Based on target type data and target attribute data, determine whether traffic targets from different video sources belong to the same physical entity; If so, all video sources corresponding to the same physical entity will be marked as associated video sources of that physical entity; The geographic coordinates in the associated video sources are connected in time stamp order to form a continuous spatiotemporal trajectory of the target, which is then stored in the historical trajectory database.

Citation Information

Patent Citations

  • Driving video recording method and system based on four-way monitoring

    CN120147987A

  • Al-based video content analysis method and system

    CN120783268A