Public place danger early warning model training method, danger early warning method and medium
Patent Information
- Application Number
- CN202512010541.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-12-29
AI Technical Summary
[0006]本发明旨在解决现有公共场所安防技术中存在的被动响应、依赖人力、信息单一、缺乏预测能力等问题,提供一种公共场所危险预警模型训练方法、危险预警方法及介质
(1)主动预警,防患于未然:本发明不再局限于对已发生事件的检测,而是通过学习历史数据中的危险前兆模式,实现对未来风险的预测,将安防模式从“被动响应”升级为“主动预防”。
Smart Images

Figure CN122067370B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a method for training a public place danger warning model, a danger warning method, and a medium. Background Technology
[0002] With the acceleration of urbanization and the increasing frequency of public activities, the security management of public places faces unprecedented challenges. Traditional security systems mainly rely on a large number of deployed surveillance cameras for video collection, with security personnel monitoring the situation manually or reviewing recorded footage afterward. However, this traditional model has many insurmountable drawbacks.
[0003] First, manual monitoring is extremely inefficient and unreliable. In large public places, there can be hundreds or even thousands of surveillance cameras, generating massive amounts of video data. Security personnel need to monitor multiple screens simultaneously, and prolonged work can easily lead to visual fatigue and distraction, resulting in a high rate of missed detections of abnormal events. Therefore, relying solely on manpower cannot achieve full-time, effective coverage of all monitored areas.
[0004] Secondly, passive monitoring cannot provide early warning. Whether it's manual monitoring or simple motion-detection-based intelligent analysis, the essence is that alarms are only triggered after a dangerous event has already occurred or is in the process of occurring. For example, when the system detects abnormal crowd gatherings, fights, falls, or other behaviors, the event has often already caused adverse effects or losses. This "after-the-fact" security model cannot nip security risks in the bud and lacks the ability to proactively prevent and intervene in advance.
[0005] Secondly, the system lacks the ability to continuously track and analyze targets across multiple cameras. Surveillance networks in public places typically consist of multiple cameras covering different areas. A dangerous act often involves the coordinated movement of multiple targets, whose trajectories may cross the fields of view of multiple cameras. Traditional single-camera analysis systems cannot construct a complete behavioral chain of a target in global space, making it difficult to understand the interaction relationships between multiple targets in complex scenarios, and thus hindering the accurate assessment of the potential danger and development trend of an event. Summary of the Invention
[0006] This invention aims to solve the problems of passive response, reliance on manpower, limited information, and lack of predictive ability in existing public place security technologies, and provides a training method for a public place danger early warning model, a danger early warning method, and a medium.
[0007] The first aspect of this invention provides a method for training a public place hazard early warning model, comprising the following steps: Acquire historical monitoring data containing confirmed hazardous events and synchronized audio; Multi-target cross-camera trajectory tracking is performed on the historical monitoring data, and audio and video features are extracted at fixed time intervals to form regional monitoring vectors. Then, the regional monitoring vectors of multiple consecutive time points are arranged in chronological order to form time-series monitoring vectors. The occurrence time of each dangerous event in the historical monitoring data is marked, and an analysis interval with a preset duration is defined that is located before the occurrence time. If the end time of the time-series monitoring vector is located within the analysis interval, it is marked as the first sample; otherwise, it is marked as the second sample. Based on the time-series monitoring vector and the corresponding sample identifier, the pre-trained model is trained to obtain a public place danger early warning model for outputting the probability of future dangerous events.
[0008] A second aspect of the present invention provides a method for early warning of danger in public places, comprising the following steps: Video streams and synchronous audio streams are acquired in real time from surveillance cameras in various monitoring areas of public places, and the video streams are preprocessed with image enhancement. Multi-target cross-camera trajectory tracking is performed on the video stream after image enhancement preprocessing, and video features of the video stream and audio features of the audio stream are extracted at fixed time intervals; The extracted video and audio features are concatenated to obtain the area monitoring vector at the current moment. The latest multiple area monitoring vectors are then arranged in chronological order to form a time-series monitoring vector. The time-series monitoring vector is input into a pre-trained public place danger early warning model, which outputs probability vectors of various dangerous events that will occur in each monitoring area of the public place in the future. Determine whether there is a probability component in the probability vector that is greater than the warning threshold; If a probability component exceeds the warning threshold, a hazard warning is triggered, generating warning information that includes the type of hazard event, the area where the hazard occurs, and a summary of key features, and security joint control is initiated.
[0009] A third aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the above-described public place hazard warning model training method or the above-described public place hazard warning method.
[0010] Compared with existing technologies, the public place danger early warning model training method proposed in this invention has the following significant advantages: (1) Proactive early warning, prevention before problems occur: This invention is no longer limited to the detection of events that have already occurred, but rather learns the danger precursor patterns in historical data to predict future risks, upgrading the security mode from "passive response" to "proactive prevention".
[0011] (2) Multimodal fusion to improve accuracy: Innovatively, video features (target attributes, behavioral status) and audio features (event sound, ambient sound) are deeply fused. The complementary information is used to significantly reduce the false alarms and false alarms that may be generated by a single modality, thereby improving the reliability and accuracy of the early warning.
[0012] (3) Global spatiotemporal analysis to understand complex scenarios: By using multi-target cross-camera trajectory tracking technology, the complete spatiotemporal trajectory of the target in the global monitoring area can be constructed, which can better understand the interaction relationship and complex behavior patterns between multiple targets, and provide a global perspective for accurate risk assessment.
[0013] (4) Enhance image quality and ensure all-weather operation: It integrates an advanced adaptive image enhancement preprocessing module, which can effectively cope with the impact of harsh environments such as rain, snow, fog, and low light on the monitoring screen, and ensure that the system can work stably and efficiently under various complex environmental conditions. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of an embodiment of the public place danger early warning model training method in the present invention; Figure 2 This is a schematic diagram of one embodiment of the public place danger early warning method in the present invention; Figure 3 This is a schematic diagram of one embodiment of the image enhancement preprocessing method in this invention. Detailed Implementation
[0015] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the public place danger early warning model training method in this invention includes: 101. Obtain historical monitoring data containing confirmed hazardous events and synchronized audio; This step is the cornerstone of the entire model training process, and the quality and richness of the data directly determine the performance of the final early warning model. In specific implementation, the historical monitoring data should meet the following requirements: First, data sources should cover a wide range of public spaces, such as, but not limited to, airport waiting halls, train station platforms and plazas, subway station transfer passages, the interiors and entrances of large shopping malls, stadium stands and surrounding areas, and city center squares. This diversity of scenarios helps the model learn the warning signs of danger in different environments, enhancing its generalization ability.
[0016] Secondly, the data should include multiple, synchronized video and audio streams. "Synchronization" here refers to strict alignment in timestamps; that is, each video frame should have a precisely corresponding audio segment. This is a prerequisite for subsequent multimodal feature fusion. Data acquisition should utilize monitoring equipment with good audio acquisition capabilities to ensure the clarity and identifiability of the audio signal.
[0017] Secondly, the data must contain a sufficient number of confirmed hazardous events. These hazardous events serve as the source of "positive example" labels for the model's learning. The types of events should be predefined, and may include, for example: fights, robberies, stampedes, sudden illness leading to collapse, explosive threats, vehicles illegally entering restricted areas, arson, etc.
[0018] Finally, to ensure the model's robustness, the historical data should also include a large number of normal monitoring periods during which no dangerous events have occurred. This data will serve as "negative examples" for the model's learning, helping it distinguish between normal states and states showing early signs of danger. Normal data should cover different time periods (e.g., daytime, nighttime), different weather conditions (e.g., sunny, rainy, snowy, foggy), and different population densities (e.g., sparse, dense, congested), to reduce the interference of environmental factors on the model's judgment.
[0019] 102. Perform multi-target cross-camera trajectory tracking on the historical monitoring data, extract audio and video features at fixed time intervals to form regional monitoring vectors, and then arrange the regional monitoring vectors of multiple consecutive time points in chronological order to form time-series monitoring vectors. The goal of this step is to transform raw, high-dimensional, unstructured audio and video data into a structured temporal vector sequence that can characterize the dynamic evolution of a scene. This process serves as a bridge connecting raw perceptual data with advanced cognitive prediction, and its granularity and comprehensiveness directly determine the richness of information the model can learn. The specific implementation methods of each detailed step (1021-1025) included in step 102 are described in detail below.
[0020] 1021: Initialize a multi-object tracking engine and load a trained object detection model and a trained object re-identification model.
[0021] The multi-target tracking engine serves as the scheduling hub and state manager for the entire process. It coordinates the operation of various sub-modules and maintains the real-time status of all tracked targets, including their identity, location, and trajectory. This engine can be customized based on existing multi-target tracking frameworks (such as DeepSORT and JDE) to meet the specific needs of this invention.
[0022] (1) Object detection model: Its task is to quickly and accurately locate all moving objects of interest in each frame of an image. This model usually adopts advanced object detection algorithms based on deep learning, such as the YOLO series, SSD, or Faster R-CNN. During training, these models use a large amount of labeled image data to enable them to identify objects of specific categories (such as "people", "cars", "bicycles") and output their precise locations in the image in the form of bounding boxes, along with a confidence score to indicate the reliability of the detection.
[0023] (2) Target Re-identification Model: Its core function is to solve the problem of identity association of targets under different time periods, different camera views, or after temporary occlusion. This model is a deep learning classification network built on backbone networks such as ResNet and Inception. Its input is a cropped image block of a target (such as a pedestrian), and the output is a fixed-length, high-dimensional deep appearance feature vector. This vector is a highly condensed and abstract representation of the target's appearance information (such as clothing color, body shape, hairstyle, and carried items). A good target re-identification model can make the feature vectors of the same target image generated under different viewpoints, lighting, and occlusion conditions very close in the vector space, while the feature vectors of different targets are far apart.
[0024] 1022: For each frame of video image in the historical monitoring data, run the target detection model to detect all moving targets and obtain the bounding box of each moving target.
[0025] This step processes the historical video data frame by frame. For each frame of each video stream in the historical monitoring data, the system inputs it into the object detection model loaded by the multi-object tracking engine. The model outputs a detection list, where each item contains information about one or more detected objects, specifically including: (1) Bounding box: Usually represented by four values, namely the coordinates of the upper left corner and the lower right corner of the target box; (2) Category label: For example, “person”, “car”; (3) Confidence score: A floating-point number between 0 and 1, representing the model’s confidence in the detection result.
[0026] The system will filter out low-quality detection results based on a preset confidence threshold (e.g., 0.5), and only retain high-confidence targets as candidates for subsequent tracking.
[0027] 1023: Using a target re-identification model, a depth appearance feature vector is calculated for each detected moving target, and combined with motion information, a unique identifier is assigned to the same moving target appearing in consecutive frames for generating single-camera trajectories.
[0028] This step involves concatenating discrete, frame-by-frame detection results into a continuous trajectory within the field of view of a single camera. This is a data association process, and its core challenge lies in handling issues such as target occlusion, deformation, rapid movement, and the appearance of new targets while old targets disappear. The specific implementation method is as follows: For each target detected in step 1022, the system crops a corresponding image patch from the original image based on its bounding box, and inputs this image patch into the target re-identification model loaded in step 1021 to obtain a depth appearance feature vector. The system utilizes motion information to assist in the association, specifically by building a motion model for each tracked target, such as a Kalman filter. This filter can predict the possible location of the target in the current frame based on the target's position and velocity in the past few frames.
[0029] For each detection result in the current frame, the system matches it with all existing tracking trajectories from the previous frame. The matching is based on the similarity of the apparent feature vectors (typically measured using cosine similarity or Euclidean distance) and the overlap between the predicted motion location and the actual detection location (e.g., measured using Intersection over Union (IoU)). A comprehensive cost function associates the current detection result with the best-matching trajectory from the previous frame and assigns them the same unique identifier. If a current detection result cannot match any existing trajectory, it is considered a newly appearing target and assigned a new ID. If an existing trajectory does not find a matching detection result within several consecutive frames, it is considered that the target has left the field of view, and its trajectory is terminated. In this way, the system generates a continuous single-camera trajectory with a unique ID for each continuously appearing target within the field of view of a single camera.
[0030] 1024: By applying cross-camera re-identification technology, the apparent feature vectors of all moving targets appearing in the fields of view of different cameras are compared, and the trajectories of single cameras belonging to the same moving target are spatiotemporally correlated to construct the complete cross-regional spatiotemporal trajectory of the moving target within the global monitoring area.
[0031] When a target disappears from the field of view of one camera and reappears in the field of view of another camera, cross-camera correlation is required. This step uses the topological relationships between cameras (such as geographical location and connection paths), the time difference between the target's disappearance and reappearance (which must be within a reasonable time range, for example, estimated based on the distance between two points and the maximum moving speed), and the similarity of the apparent feature vectors generated by the target re-identification model to determine whether the trajectories under different cameras belong to the same physical target.
[0032] For example, if trajectory 1 of camera A disappears at time T, and trajectory 2 of camera B appears at time T+Δt, and the two cameras are physically adjacent, Δt is within a reasonable range, and the appearance of the end target of trajectory 1 is highly similar to the appearance of the beginning target of trajectory 2, then it can be determined that these two trajectories belong to the same target, and thus they can be connected to form a complete spatiotemporal trajectory spanning multiple monitoring areas.
[0033] This step integrates isolated, local single-camera trajectories into global, continuous cross-camera trajectories, which is crucial for achieving global situational awareness. Its core lies in solving the correlation problem of "a target disappearing at point A and then reappearing at point B." The specific implementation process involves constraints and matching at multiple levels: (1) Spatiotemporal Constraint Filtering: This is the first layer of filtering, using common sense from the physical world to eliminate a large number of impossible associations. The system needs to know in advance the geographical location and field of view coverage of each camera in the monitoring network. When a trajectory disappears at the edge of the field of view of camera A, the system will find camera B that may continue in spatiotemporal space according to the camera network topology. For example, only if camera B is physically adjacent to camera A, and the distance from the point of disappearance of A to the point of appearance of B is reachable within a reasonable time difference (estimated based on the average walking or running speed of a person), then the trajectory appearing in camera B is considered a candidate matching trajectory.
[0034] (2) Appearance feature matching: For candidate trajectory pairs selected through spatiotemporal constraints, the system extracts the target image patches at the times of their disappearance and appearance, and uses the target re-identification model to calculate their depth appearance feature vectors. Then, the similarity between the two vectors is calculated. If the similarity exceeds a preset threshold, they are considered to be the same target visually.
[0035] (3) Global Optimization and Decision-Making: In complex scenarios, there may be multiple possible associated paths, forming a complex association graph. To obtain the globally optimal solution, a graph optimization algorithm can be used, where each single camera trajectory is treated as a node in the graph, and the edges between nodes represent possible association relationships (the weights are determined by spatiotemporal and apparent similarity). Then, by finding the optimal path or the maximum weight match in the graph, the final cross-camera association result is determined.
[0036] Through this series of complex operations, the system can seamlessly connect all trajectory segments captured by all cameras along the path a target takes from entering the entire monitored area to leaving, forming a complete, cross-regional spatiotemporal trajectory. This trajectory includes the target's complete path, speed changes, and detailed appearance information from each camera throughout the process.
[0037] 1025: Take multimodal feature snapshots of each monitored area at a fixed time interval.
[0038] This step transforms the continuous spatiotemporal trajectory and raw audio / video signals into structured, quantifiable feature descriptions at discrete time points. The fixed time interval can be set according to actual needs, such as once per second or every two seconds. At each time point, the system takes a "snapshot" of each monitored area, extracting key information. Specifically, this includes: The system iterates through all tracked moving targets within the monitored area. For each moving target, a pre-trained attribute recognition network is used to extract video features from the corresponding image region. Simultaneously, a pre-trained audio classification network is used to extract audio features from the synchronized audio signal within the current time window.
[0039] For a more detailed explanation, refer to the refinement steps 10251-10253 of step 1025: 10251: When performing video feature extraction, for each tracked moving target, the basic attributes and behavioral state attributes are extracted from the image region corresponding to the moving target using a trained attribute recognition network at fixed time intervals.
[0040] An attribute recognition network is a deep learning model that can consist of one or more neural networks specifically designed for attribute recognition, typically using convolutional neural networks, and is trained under supervised supervision for a specific attribute recognition task. The extracted attributes fall into two main categories: basic attributes and behavioral / state attributes.
[0041] (1) For personnel targets, their basic attributes describe their relatively static, inherent, or unchanging visual characteristics, including: 1) Gender: male or female; 2) Age group: can be divided into children, youth, middle-aged, elderly, etc.; 3) Height: can be divided into multiple categories, such as tall, medium, and short; Body type: can be divided into multiple categories, such as thin, standard, and fat.
[0042] (2) The behavioral state attributes of personnel targets are more dynamic, reflecting the immediate behavior and state of the targets, including: 1) Whether carrying luggage: such as suitcases, backpacks, etc.; 2) Whether carrying suspected knives, such as long, thin, reflective, or other sharp objects; 3) Whether masked: such as being largely obscured by masks, hats, sunglasses, etc.; 4) Number of companions: by analyzing the trajectories of other targets within a certain space and time window around the target, determine whether the target is walking alone or with a group of people, and count the number of companions; 5) Whether running: by analyzing the speed and amplitude of movement at key points of the target's movement to determine whether the target is running; 6) Whether abnormal physical movements are observed: such as waving arms, aggressive postures, falling, distress gestures, etc.
[0043] (3) For vehicle targets, their basic attributes include: 1) License plate number; 2) Vehicle type: such as sedan, bus, truck, SUV, etc.
[0044] (4) The behavioral state attributes of the vehicle target include: 1) Whether the vehicle has entered a restricted area: Determine if there is a vehicle in the restricted area; 2) Whether the warning lights are on: Determine if the warning lights are on by detecting the flashing lights on the top of the vehicle or in a specific location.
[0045] 10252: When performing audio feature extraction, a trained audio classification network is used to extract event sound features and environmental sound features for synchronous audio signals at fixed time intervals.
[0046] Audio analysis is also performed at fixed time intervals, with the analysis window being a short period of time preceding the current moment, such as the first two seconds of audio. The extracted features are also divided into two categories: (1) Event sound characteristics refer to sounds that are closely related to a specific dangerous event and have clear semantics, including: 1) Detection of screams; 2) Cries for help: identified by recognizing keywords such as "help"; 3) Sounds of heated arguments: determined by analyzing the audio's spectral characteristics (such as energy, zero-crossing rate, Mel-frequency cepstral coefficients, MFCCs) and prosodic features to determine if it features multiple people speaking loudly and rapidly alternating; 4) Sounds of breaking glass: identified by recognizing the unique transient, high-frequency impact acoustic patterns of breaking glass; 5) Sounds of vehicle horns: identified by recognizing the specific frequency and time-domain waveform of vehicle horns.
[0047] (2) Ambient sound characteristics describe the overall sound environment of the scene. They are not specific to any particular event, but they reflect the overall atmosphere of the scene, including: 1) Background volume in decibels: A numerical characteristic obtained by calculating the root mean square energy of the audio signal and converting it to decibels, reflecting the overall volume level; 2) Crowd noise level: This can be quantified by analyzing the spectral complexity and energy distribution uniformity of the audio signal. For example, a noisy environment typically has a wider spectrum and a more irregular energy distribution, while a quiet environment is the opposite.
[0048] 10253: The extracted categorical attributes are vectorized using one-hot encoding, the extracted numerical attributes are normalized and vectorized, and all vectorized video and audio features are concatenated to form a high-dimensional feature vector.
[0049] (1) One-hot encoding: For discrete category attributes such as gender and whether or not luggage is carried, one-hot encoding is used. For example, "gender: male" can be encoded as a two-dimensional vector [1, 0], and "gender: female" is [0, 1]. "whether or not luggage is carried: yes" can be encoded as [1, 0], and "no" is [0, 1].
[0050] (2) Normalization: For continuous numerical attributes such as height and background volume in decibels, normalization is required, such as linearly scaling their values to the range of [0, 1] to eliminate the differences in the dimensions of different attributes and prevent attributes with larger values from dominating the model training.
[0051] After the above processing, each attribute of each target becomes a numerical vector. Finally, the system concatenates the video feature vectors of all tracked targets within a monitoring area (one vector per target) and the global audio feature vector of that area in a predetermined order, forming a very long, single, high-dimensional vector. This vector is the area monitoring vector. Mathematically, it completely and structurally describes the static attributes, dynamic behaviors, and overall sound state of the environment of all key targets within a specific monitoring area at a specific timestamp.
[0052] 1026: All extracted video and audio features are vectorized, and the video feature vectors of all moving targets in the same monitoring area at the same time are concatenated with the global audio feature vector to form a high-dimensional regional monitoring vector.
[0053] The regional monitoring vector is an information aggregate, a unified representation of micro-level individual characteristics (each target) and macro-level environmental characteristics (global audio) within a specific spatiotemporal slice. It provides an extremely rich input for subsequent time-series models.
[0054] 1027: Arrange the regional monitoring vectors of N consecutive time points in chronological order to form a time-series monitoring vector of length N, and use it as a basic training sample for the pre-trained model.
[0055] A single area monitoring vector is merely a static snapshot, unable to capture the evolution of behavior and the development trend of events. To learn dynamic patterns, a time dimension must be introduced. A time window length N is set, for example, N=10, representing 10 consecutive time points (i.e., 10 seconds, if the interval is 1 second) of area monitoring vectors. Stacking these 10 vectors in chronological order forms a two-dimensional matrix (N x D, where D is the dimension of a single area monitoring vector), or it can be flattened into a longer one-dimensional vector, which is the temporal monitoring vector. Each such temporal monitoring vector constitutes a basic input sample for a pre-trained model. It contains the dynamic changes of all relevant information within the monitored area from time T-N+1 to time T, such as how the target moves, how crowds gather, and how the sound escalates, enabling the model to learn the evolutionary patterns before a dangerous event occurs.
[0056] 103. Mark the occurrence time of each dangerous event in the historical monitoring data, and define an analysis interval that is located before the occurrence time and has a preset duration. If the end time of the time-series monitoring vector is located within the analysis interval, it is marked as the first sample; otherwise, it is marked as the second sample. This step is a crucial bridge connecting the raw feature data with supervised model training. Its core task is to assign explicit semantic labels to the large number of unlabeled time-series monitoring vectors generated in step 102, thereby constructing a high-quality training dataset for the model to learn from. This process essentially answers a core question: "What is the state of the system (i.e., the time-series monitoring vectors) before a dangerous event actually occurs?" By accurately defining and labeling these pre-event states and the state before the event, the model can learn to distinguish them and thus achieve prediction. Step 103 can be further broken down into the following sub-steps: 1031: Deploy a pre-trained hazard event recognition model, wherein the hazard event recognition model is a multimodal classifier based on a convolutional neural network-long short-term memory network.
[0057] Faced with massive amounts of historical monitoring data, relying entirely on manual frame-by-frame annotation of hazardous events is an extremely time-consuming, labor-intensive, and error-prone task. To improve annotation efficiency and consistency, this step first introduces an auxiliary automated tool. This hazardous event identification model is a deep learning model specifically designed to identify hazardous events that have already occurred. It employs a multimodal architecture of Convolutional Neural Network-Long Short-Term Memory Network (CNN-LSTM): (1) Convolutional Neural Network (CNN) part: responsible for extracting local features in the spatial and frequency domains from each frame of video image and the corresponding audio spectrogram. For video, CNN can learn visual patterns such as edges, textures, and object parts; for audio, CNN can learn the acoustic features of different sound events.
[0058] (2) Long Short-Term Memory (LSTM) part: responsible for processing the feature sequences extracted by CNN. As a special type of recurrent neural network, LSTM can effectively capture long-distance dependencies in time series, for example, identifying a “fight” event from a series of continuous arguing sounds and intense actions.
[0059] 1032: Using the aforementioned hazardous event identification model, the historical monitoring data is scanned and analyzed frame by frame, and the specific time point, specific monitoring area, and event category of each type of hazardous event are marked.
[0060] The system inputs historical monitoring data (synchronized video and audio streams) into the hazard event identification model deployed in step 1031. The hazard event identification model scans and analyzes the data frame by frame (or in small segments) using a sliding window approach. When the model detects that the multimodal features at a certain point in time highly match the pattern of a certain type of hazard event, it outputs a label. This label information contains at least three key elements: (1) Specific time point: The precise timestamp at which the event was identified, such as 2023-10-27 15:32:18. (2) Specific monitoring area where the event occurred: The camera ID or area number where the event occurred, such as Cam_03_Area_B. (3) Category of the event: Such as fighting, robbery, stampede, etc.
[0061] 1033: For each marked hazardous event, define a hazardous event analysis interval located before the occurrence time corresponding to the hazardous event. The analysis interval includes the start time, end time, and duration.
[0062] This step shifts the focus from when an event occurs to before it occurs.
[0063] (1) Analysis interval: For each hazardous event identified in step 1032, its occurrence time is T_event. The analysis interval is a time window [T_start, T_end].
[0064] (2) End time: The end time T_end of the event analysis interval is set to the event occurrence time T_event. That is, we are concerned with the state at the last moment before the event occurs.
[0065] (3) Start Time: The start time T_start is calculated from T_event - Duration. Here, Duration is a preset duration, such as 30 seconds, 60 seconds, or longer. This duration is a key hyperparameter, and its value needs to be set empirically according to the characteristics of different types of dangerous events. For example, some events triggered by emotional excitement (such as fights) may have a long warning period, while some sudden events (such as explosions) may have a very short warning period.
[0066] By defining this pre-analysis interval, clear rules are set for subsequent sample labeling: any state data, as long as its timestamp falls within this window, is considered a precursor to the event.
[0067] 1034: Create an empty training dataset and iterate through all time-series monitoring vector samples.
[0068] The system initializes a data structure in memory or storage to store the final training sample pairs (input data, labels). Then, the system begins processing all the time-series monitoring vector samples generated in step 102 one by one. Each sample represents a sequence of scene state snapshots over a specific time period (e.g., from T-9 seconds to T seconds).
[0069] 1035: For each time series monitoring vector sample, check whether the time series monitoring vector sample falls within any predefined analysis interval at the end time.
[0070] For each time-series monitoring vector sample being processed, there is a defined end time T_end_sample. The system checks whether this T_end_sample falls within the analysis interval [T_start_event, T_end_event] defined in step 1033 for any hazardous event. This is a simple time inclusion relationship determination. Since historical data may contain multiple hazardous events, the end time of a sample may fall within multiple analysis intervals, or it may not fall within any of them.
[0071] 1036: If the time series monitoring vector sample falls within the analysis interval of the dangerous event at the end time, then create a first sample label of an M-dimensional vector for the time series monitoring vector sample, where M is the total number of dangerous event categories, and set the component corresponding to the dangerous event in the first sample label to 1, and set the other components to 0.
[0072] This step defines the labeling method for positive samples (first samples).
[0073] Triggering condition: If the judgment result of step 1035 is "yes", that is, T_start_event<= T_end_sample<= T_end_event is true for a certain event.
[0074] Label vector: Create an M-dimensional label vector y, where M is the total number of all predefined dangerous event categories in the system (e.g., if the focus is on fights, robberies, stampedes, falls, and intrusions, then M=5).
[0075] Labeling Rules: This is a multi-label classification problem. If the sample falls within the analysis interval of the "fighting" event, then the component corresponding to "fighting" in the label vector y is set to 1. If the sample also falls within the analysis interval of the "robbery" event (e.g., the precursors of the two events are intertwined), then the component corresponding to "robbery" is also set to 1. Components corresponding to all other unrelated event categories remain 0. For example, if M=5, and the order is [fighting, robbery, stampede, falling, breaking in], a sample falling into the "robbery" interval would be labeled [0, 1, 0, 0, 0].
[0076] 1037: If the time series monitoring vector sample does not fall within the analysis interval of any dangerous event at the end time, then create a second sample label for the time series monitoring vector sample, which is an M-dimensional vector with all components being 0.
[0077] This step defines the labeling method for negative samples (second samples).
[0078] Triggering condition: If the judgment result of step 1035 is "no", that is, the end time T_end_sample of the sample does not fall into any defined analysis interval.
[0079] Label vector: Similarly, create an M-dimensional label vector y, but set all components of this vector to 0, i.e., [0, 0, 0, 0, 0]. This all-zero vector intuitively represents "no dangerous events are about to occur in the next time period", that is, a normal or safe state.
[0080] 1038: Pair all the constructed time-series monitoring vector samples with their corresponding sample identifiers according to time points to form a complete model training dataset.
[0081] For each time-series monitoring vector sample X_i, after processing in steps 1035-1037, a unique corresponding M-dimensional label vector y_i is obtained. The system pairs them as (X_i, y_i). Collecting all sample pairs constitutes a complete, structured model training dataset {(X1, y1), (X2, y2),..., (X_i, y_i)} with precise labels. n , y n This dataset can now be directly used in the model training process of step 104, clearly telling the model what probability distribution y should be output when given a time-series state X as input.
[0082] Through the steps described above, this embodiment transforms a vague prediction requirement into a well-defined classification task that can be solved by a machine learning model. This not only provides the model with the answers (labels) to learn, but more importantly, it defines the problem itself, namely, identifying specific state patterns before a dangerous event occurs, which is the foundation for achieving proactive early warning.
[0083] 104. Based on the time-series monitoring vector and the corresponding sample identifier, the pre-trained model is trained to obtain a public place danger early warning model for outputting the probability of future dangerous events.
[0084] Through steps 101-103, a high-quality training dataset with accurate labels was obtained. The goal of this step is to use this dataset to train a deep neural network model through a systematic and iterative optimization process, enabling it to accurately infer the probability of various dangerous events occurring in the future from the input time-series monitoring vectors. This process essentially involves finding a complex functional mapping relationship to map the high-dimensional time-series feature space to a low-dimensional danger probability space. Step 104 can be further broken down into the following sub-steps: 1041: Select a Long Short-Term Memory (LSTM) network as the basic architecture of the pre-trained model and configure the network structure parameters. Connect a fully connected layer after the output layer of the LSM network and use a preset activation function to make the pre-trained model output a multi-dimensional probability vector, where the value of each component represents the probability of a corresponding category of dangerous event occurring in the future.
[0085] Infrastructure Selection: This embodiment chooses Long Short-Term Memory (LSTM) as the core because LSTM is a special type of recurrent neural network that effectively solves the gradient vanishing or exploding problems that traditional RNNs may encounter when processing long sequences through its ingenious "gating" structure (input gate, forget gate, output gate). This allows LSTM to capture long-term dependencies spanning multiple time steps in time-series data. For hazard warning, the precursors to an event may begin to appear tens of seconds or even minutes in advance, making this long-term memory capability of LSTM crucial.
[0086] Network architecture parameter configuration: Configuring network architecture parameters involves fine-tuning the model's complexity. This includes: determining the number of LSTM layers (e.g., using a stacked 2 or 3-layer LSTM to learn more complex hierarchical features); setting the number of hidden units in each LSTM layer (more hidden units result in a larger model capacity, enabling the learning of more complex patterns, but also making it more prone to overfitting); and setting the dropout layer ratio (randomly setting the output of a subset of neurons to zero during training to enhance the model's generalization ability and prevent overfitting).
[0087] Output layer design: A fully connected layer is connected at the top of the LSTM layer. The function of this fully connected layer is to linearly transform the high-level temporal features extracted by the LSTM, mapping their dimensions to the same dimension as the number of hazard event categories M.
[0088] Activation Function Selection: A preset activation function is applied after the fully connected layer. In the multi-label classification scenario of this embodiment, when a dangerous event occurs, it may be accompanied by other types of risks (for example, a fight may be accompanied by screams and the sound of breaking glass), and the categories are not mutually exclusive. Therefore, the Sigmoid function is preferred. The Sigmoid function can compress any real value of the fully connected layer output into the interval (0, 1), so that each component of the output vector can be independently interpreted as the probability of the corresponding category of dangerous event occurring. If the problem is defined as multi-class mutually exclusive classification (i.e., only one event can occur at a time), the Softmax function can be used to ensure that the sum of all output components is 1.
[0089] 1042: Divide the prepared training dataset into a training set and a validation set according to a preset ratio.
[0090] 1043: Set the binary cross-entropy as the loss function of the model, select the preset optimizer as the optimization algorithm, and set the initial learning rate and weight decay coefficient.
[0091] Loss Function Setting: The loss function quantifies the difference between the model's predicted value and the true label. For the multi-label classification problem in this embodiment, binary cross-entropy is preferably used as the loss function. For each sample, the total loss is the sum of its binary cross-entropy losses across the M classes. Specifically, for the i-th class, if the true label is 1 (event occurred) and the model predicts a probability of p, then the loss for that class is -log(p); if the true label is 0 (event did not occur) and the model predicts a probability of p, then the loss is -log(1-p). The model's goal is to minimize the total loss across the entire training set by adjusting the weights.
[0092] Optimizer Selection: The optimizer is the specific algorithm that minimizes the loss. It updates the model's weights based on the gradient calculated from the loss function. Commonly used optimizers include stochastic gradient descent, momentum-based SGD, AdaGrad, RMSProp, and Adam. The Adam optimizer is preferred because it combines the advantages of momentum and adaptive learning rates, typically performs well, and is less sensitive to hyperparameters.
[0093] Initial learning rate: This is the step size used by the optimizer at the start of training. The learning rate is one of the most important hyperparameters during training. If the learning rate is too high, the loss function may oscillate around the optimal value or even diverge; if the learning rate is too low, the training speed will be very slow. A suitable initial value is usually chosen experimentally.
[0094] Weight decay: This is a regularization technique that encourages the model to learn smaller weight values by adding an L2 norm penalty term to the loss function for the model weights. This helps simplify the model, prevents it from overfitting to noise in the training data, and thus improves generalization ability.
[0095] 1044: Start the training process, input the training set data into the pre-trained model in batches, perform forward propagation calculation, calculate the loss between the predicted probability and the true probability, then calculate the gradient through the backpropagation algorithm, and update the network weights using the preset optimizer.
[0096] Forward propagation: For a batch of data, the time-series monitoring vector is input into the model. The data flows sequentially through the LSTM layer and the fully connected layer, and finally obtains a predicted probability vector through the activation function.
[0097] Loss calculation: Using the binary cross-entropy loss function set in step 1043, the predicted probability vector output by the model is compared with the true label vector corresponding to the batch of data to calculate the average loss of the batch.
[0098] Backpropagation: This is the core algorithm of deep learning. Starting with the loss function, it uses the chain rule from calculus to calculate the gradient of the loss function with respect to each parameter (weight and bias) in the model, layer by layer. The gradient indicates the direction and magnitude by which each parameter should be adjusted to reduce the loss.
[0099] Weight update: The optimizer adjusts all parameters of the model according to its specific update rules (such as Adam's update formula) based on the calculated gradient and the set learning rate. After one update, one iteration is completed.
[0100] 1045: After each training cycle, use the validation set to evaluate the model's performance and monitor the loss value and accuracy on the validation set.
[0101] A training cycle refers to the time when the model has been run through all the samples in the training set once.
[0102] Performance Evaluation: At the end of each cycle, the model switches to evaluation mode. In this mode, all data from the validation set (either in batches or in large batches) is input into the model for forward propagation to obtain prediction results. Then, evaluation metrics such as total loss and accuracy on the validation set are calculated. Accuracy can be defined as the proportion of correctly predicted labels to the total number of labels, or more precise metrics such as the F1 score can be used, especially when the data is imbalanced.
[0103] Monitoring metrics: Typically, these are plotted as curves showing the changes in training loss and validation loss over time. Ideally, both should decrease as training progresses. If the training loss continues to decrease, but the validation loss begins to rise or stagnates, this is usually a clear signal that the model is starting to overfit.
[0104] 1046: If the loss value on the validation set no longer decreases over multiple consecutive training periods, terminate the model training, save the optimal model parameters, and obtain the trained public place danger early warning model.
[0105] During training, the model state is continuously recorded when the validation set loss reaches its minimum value. If the validation set loss does not reach a new minimum value within the next 10 consecutive epochs, training is terminated early. The system does not save the model from the last epoch; instead, it saves the model parameters from the epoch that performed best on the validation set (i.e., had the lowest validation loss). These saved parameters constitute the final public place hazard warning model. This model can then be deployed in a real-world warning system for analyzing and predicting real-time data.
[0106] Please see Figure 2 One embodiment of the public place danger early warning method in this invention includes: 201. The surveillance cameras in each monitored area acquire video streams and synchronous audio streams in real time, and perform image enhancement preprocessing on the video streams; Considering that the complex environment of public places can significantly affect the accuracy of hazard warnings, image enhancement preprocessing of surveillance videos has been added. This aims to overcome the impact of harsh environments on video analysis and provide high-quality, high-definition, and time-stable visual data for subsequent real-time feature extraction and hazard warnings.
[0107] 202. Perform multi-target cross-camera trajectory tracking on the video stream after image enhancement preprocessing, and extract video features of the video stream and audio features of the audio stream at fixed time intervals; Step 202 is technically identical to step 102 in the model training method, the difference being that it processes real-time data streams. In practice, the system processes multiple real-time video and audio streams in parallel. The multi-object tracking engine, object detection model, object re-identification model, attribute recognition network, and audio classification network are all prepared during the training phase. The system executes this step in real-time to continuously output the video and audio features of each monitored area at each fixed time interval.
[0108] 203. The extracted video and audio features are concatenated to obtain the area monitoring vector at the current moment, and the latest multiple area monitoring vectors are arranged in chronological order to form a time-series monitoring vector; This step is consistent with steps 1026 and 1027 in the training method. The system will vectorize and concatenate the features extracted in real time in step 202 according to preset rules to form the regional monitoring vector at the current moment. Simultaneously, the system maintains a sliding time window buffer to store the regional monitoring vectors of the most recent N time points. Whenever a new regional monitoring vector is generated, it is added to the buffer, and the oldest one is removed. In this way, the buffer always maintains the latest N vectors, which, arranged in chronological order, constitute the real-time temporal monitoring vector. This real-time temporal monitoring vector is a snapshot input to the early warning model, reflecting the dynamic evolution of the monitored scene within a short period before the current moment.
[0109] 204. Input the time-series monitoring vector into the pre-trained public place danger early warning model, and output the probability vector of various dangerous events that will occur in each monitoring area of the public place in the future; This is the core of the early warning decision-making process. The real-time time-series monitoring vector generated in step 203 is fed as input into the public place danger early warning model trained in the first part. After one forward propagation calculation, the model will quickly output an M-dimensional probability vector. For example, the output might be [0.05, 0.82, 0.12, 0.03, 0.01]. This vector indicates that, based on current and past observations, it is predicted that in the near future, the probability of a "fight" occurring in the monitored area is 5%, the probability of a "robbery" is 82%, the probability of a "stampede" is 12%, the probability of a "falling to the ground" is 3%, and the probability of an "intrusion" is 1%. This probability vector is the direct basis for the system to make the next judgment.
[0110] 205. Determine whether there is a probability component in the probability vector that is greater than the warning threshold; 206. If there is a probability component greater than the warning threshold, a danger warning will be triggered, and a warning message containing the type of dangerous event, the area where the danger occurred, and a summary of key features will be generated, and security joint control will be initiated.
[0111] To translate the model's probability output into specific early warning decisions, an early warning threshold needs to be set. This threshold is a value between 0 and 1, such as 0.7 or 0.8. The system checks each component in the probability vector output in step 204. If at least one component's value exceeds the preset early warning threshold—for example, the "robbery" probability of 0.82 in the above example exceeds the threshold of 0.8—the system determines that there is a high risk. If all components are below the threshold, the current state is considered normal, no early warning is triggered, and the system continues with the next round of real-time monitoring.
[0112] Once step 205 determines that there is a high risk, the system will immediately trigger a series of response actions.
[0113] First, generate an early warning message. This message is for security personnel and needs to be clear and intuitive. It should include at least: (1) Dangerous event type: that is, the event category corresponding to the component in the probability vector that exceeds the threshold, such as "robbery".
[0114] (2) Dangerous area: that is, the monitoring area that generates the high-risk time-series monitoring vector, such as "the entrance camera 3 of Area A on the second floor of the train station".
[0115] (3) Key Feature Summary: This is an interpretable demonstration of the basis for the model's judgment. The system can trace back the original features that constitute the time-series vector and present them in the form of natural language or a list of labels. For example: "Two people were detected running at high speed, one of whom was suspected of carrying a knife. Simultaneous audio detected loud arguing and screaming, and the background volume decibel level was abnormally high." Such a summary can help security personnel quickly understand the situation and verify the authenticity of the warning.
[0116] Secondly, initiate security control linkage. The system can automatically execute a series of security measures through preset interfaces, such as: (1) Automatically pop up the real-time image of the relevant monitoring area on the large screen of the monitoring center and highlight the relevant target; (2) Send early warning information and on-site images to the mobile terminal or walkie-talkie of the designated security personnel; (3) Automatically link the nearby broadcasting system to issue voice warnings; (4) In some high-risk scenarios, it can automatically lock the access control of the relevant area and even notify the relevant departments.
[0117] In traditional video enhancement methods, enhancement algorithms (such as gamma correction and histogram equalization) are typically applied globally or in a fixed manner to the entire image, much like applying a uniform base color to a painting. This one-size-fits-all approach ignores the inherent differences in image content and fails to provide personalized solutions. The pixel-level enhancement scheme diagram proposed in this invention transforms image enhancement from a signal processing problem into a cognitive decision-making and planning problem. The pixel-level enhancement scheme diagram is not an image, but a multi-dimensional data structure with the same spatial resolution (i.e., the same height and width) as the original video frame; it can be called an instruction map or control blueprint. Each pixel position (i, j) in this diagram stores one or more specific parameters that precisely define what enhancement operation should be performed on the corresponding pixel position (i, j) in the original image, the strength of the operation, and which specific algorithm submodule should be used. The process of generating this instruction map is a simulation of the thought process of human experts performing image post-processing: First, observe and understand the environment of the entire scene (is it night or rain?); second, identify different objects in the scene (where are the people, the cars, and the sky?); third, focus on the most important details (are the faces clear? Are the license plates legible?); finally, synthesize all the information to customize the most suitable processing solution for each part of the scene. This cognitive and decision-making process, from macro to micro and from global to local, is encoded into a computable and automated workflow, ultimately materializing into this pixel-level enhancement solution map.
[0118] Please see Figure 3 One embodiment of the image enhancement preprocessing method in this invention includes: 301. Perform environmental degradation classification, semantic content segmentation, and key object detection sequentially on each original video frame of the video stream to obtain scene cognition and analysis results; Traditional image enhancement methods are typically global and one-size-fits-all, such as histogram equalization or simple dehazing of the entire image. However, these methods lack an understanding of the scene content, often resulting in over-enhancement of some areas while under-enhancement of others, or even artifacts. This invention proposes an improved image enhancement preprocessing solution. This solution first understands the image content, then tailors an enhancement scheme for each video frame pixel, and finally executes it precisely through a deep learning network, ensuring the temporal coherence of the video stream. Step 301 can be further broken down into the following sub-steps: 3011: Input each original video frame of the video stream into a trained environment classification network in sequence, and output an environment degradation label representing the main degradation type of the current scene.
[0119] Environmental degradation is the primary factor affecting the quality of surveillance images. This sub-step aims to quickly and accurately identify the main problems in the current image. Environmental degradation labels can be one or more categories, such as, but not limited to: normal lighting, low light / night, rain, snow, fog, haze, shadows, motion blur, etc. The environment classification network is a specially trained convolutional neural network (CNN) whose task is not to identify objects in the image, but to determine the overall imaging conditions of the image. This label is the primary basis for the selection of subsequent enhancement strategies, because different degradation types require fundamentally different processing algorithms. For example, for "fog," a dehazing algorithm is needed; for "low light," brightness enhancement and contrast enhancement are needed; for "rain," rain line detection and removal may be required. This step sets the overall tone for the entire enhancement process.
[0120] 3012: Input each original video frame of the video stream into a trained real-time semantic segmentation network in sequence to generate a semantic segmentation map that identifies different semantic regions.
[0121] If environmental classification is about determining "what the weather is like," then semantic segmentation is about understanding "what's in the picture." Semantic segmentation networks are advanced deep learning models that assign a category label to each pixel in an image, generating a colored or numbered "semantic segmentation map" of the same size as the original image. For example, all pixels belonging to "sky" are labeled with one color, all pixels belonging to "road" are labeled with another color, and "buildings," "vegetation," "pedestrians," and "vehicles" each have their own specific labels. This semantic segmentation map provides extremely valuable spatial contextual information. Its importance lies in the fact that the expectations and tolerances for image enhancement vary greatly depending on the semantic content of the region. For example, it's desirable to enhance the clarity of the "face" region for easier recognition, but over-sharpening that distorts skin texture must be avoided; for the "sky" region, excessive processing may be undesirable to prevent color distortion; while for the "road" or "ground" region, strong contrast enhancement can be applied to highlight details. Semantic segmentation maps make "regional differentiation" enhancement possible.
[0122] 3013: Input each original video frame of the video stream into a trained key object detection model in sequence, and output a set of bounding box information containing the location and category of the key object.
[0123] This step involves the precise localization of specific important objects, complementing and refining semantic segmentation. Key objects typically refer to those requiring special treatment, primarily including faces and license plate numbers. Key object detection models (such as those based on YOLO or SSD) output a series of bounding boxes, each precisely defining the location of a key object and its category. This step serves two purposes: first, privacy protection—when enhancing images, special, conservative enhancement strategies can be applied to "face" and "license plate" regions to avoid over-processing that could lead to the leakage of personal information or damage to critical data; second, information preservation—these regions often contain crucial identification information, and enhancement processing must be done without destroying this information. The bounding box information provides the precise coordinates of "no-go zones" or "special processing areas" for subsequent pixel-level enhancement strategies.
[0124] Through the three sub-steps described in step 301, the system gains a profound understanding of each original image frame on three levels: the overall environmental degradation type, pixel-level semantic content distribution, and the precise location of key objects. These three elements together constitute a comprehensive understanding and analysis of the scene, forming the foundation for subsequent intelligent and adaptive enhancement.
[0125] 302. Based on the scene cognition and analysis results, generate a specified enhancement algorithm for each pixel of the original video frame and configure the corresponding enhancement scheme diagram; This step utilizes the rich analytical results obtained in step 301 to determine the most suitable enhancement scheme for each pixel in the image. Refer to the refinement steps 3021-3023 of step 302: 3021: Pre-establish an enhancement strategy library containing various basic enhancement algorithms and their different parameter configurations.
[0126] To generate diverse enhancement schemes for each pixel, the system requires a library of enhancement strategies, pre-loaded with various classic or deep learning-based image enhancement algorithms, each with adjustable parameters. For example, the enhancement strategy library could include: 1) Dehazing algorithms: such as algorithms based on dark channel priors, with adjustable parameters including atmospheric scattering coefficient, transmittance threshold, etc.; 2) Deraining algorithms: such as algorithms based on rain line detection and sparse coding, with adjustable parameters including rain line length, direction, density, etc.; 3) Low-light enhancement algorithms: such as algorithms based on Retinex theory or gamma correction, with adjustable parameters including gain, offset, gamma value, etc.; 4) Denoising algorithms: such as bilateral filtering or nonlocal mean denoising, with adjustable parameters including filter kernel size, smoothing parameters, etc.; 5) Sharpening algorithms: such as Unsharp Masking, with adjustable parameters including sharpening intensity, radius, etc.; 6) Contrast enhancement algorithms: such as Adaptive Histogram Equalization (CLAHE), with adjustable parameters including contrast limit, grid size, etc.
[0127] The policy library is not a simple list of algorithms, but a carefully designed, hierarchical database. Its structure contains at least three main dimensions: environment degradation type, semantic content category, and key object type.
[0128] First Dimension: Types of Environmental Degradation This dimension categorizes different severe environments. For example: {Normal, Low Illumination, Backlight, Rainy Day - Slight, Rainy Day - Medium, Rainy Day - Heavy, Snowy Day - Slight, Snowy Day - Heavy, Foggy Day - Light, Foggy Day - Heavy, Dust Storm, Mixed Degradation}.
[0129] Each degradation type is associated with a set of basic global enhancement algorithms and their default parameters. For example, for "low illumination", the default is "Retinex algorithm A, parameter set {gamma coefficient = 1.2, intensity = 0.7}"; for "rainy day - medium", the default is "rain removal model B, parameter set {network depth = 5 layers, filter intensity = 0.8}".
[0130] Second dimension: Semantic content category This dimension is based on the semantic segmentation results of the image. For example: {people, vehicles, roads, buildings, sky, vegetation, water, others}.
[0131] Each semantic category defines how the basic enhancement strategy for the environment should be adjusted within that category. For example, for the "People" category, the rule might be: "In low-light environments, increase the intensity of the Retinex algorithm by 15% and enable the face bounding box color distortion protection module to prevent face overexposure distortion." For the "Roads" category, the rule might be: "In rainy environments, enhance the intensity of the rain removal model and enable the reflection suppression algorithm."
[0132] Third Dimension: Key Object Types This dimension has the highest priority, focusing on the most critical details in security tasks. Examples include: {faces, hands, license plates, suspicious items}.
[0133] Each key object type is associated with a specialized, refined enhancement scheme. These schemes often do not use general enhancement algorithms, but rather optimized algorithms tailored to that object. For example, for "face," the associated scheme might be "a face-specific low-light enhancement network C, which is trained to specifically enhance the preservation of facial features and the accurate reproduction of colors in the face bounding box region"; for "license plate," the associated scheme might be "a high dynamic range (HDR) synthesis algorithm D, used to simultaneously see characters in bright and dark areas in backlight."
[0134] Furthermore, each rule or scheme in the policy library must ultimately be transformed into parameters that the computer can understand. These parameters constitute the specific content of each pixel vector in the pixel-level enhancement scheme graph. A typical parameter vector might contain the following elements: [Algorithm selection ID, Enhancement strength_1, Enhancement strength_2, Special switch_1, Special switch_2, ...] (1) Algorithm Selection ID: An integer or code used to select one from multiple enhancement algorithm submodules. For example, 0 represents no operation, 1 represents Retinex algorithm A, 2 represents rain removal model B, and 3 represents face-specific network C.
[0135] (2) Enhancement strength: One or more floating-point numbers used to control the strength of the selected algorithm. For example, gamma coefficients, filter weights, scaling factors of network output layers, etc.
[0136] (3) Special switch: one or more binary (0 or 1) or enumeration values to enable or disable specific functional modules, such as "face frame area color distortion protection", "reflection suppression", "edge sharpening", etc.
[0137] (4) Through this highly structured and parameterized design, the strategy library provides a solid foundation for generating precise and flexible pixel-level instructions.
[0138] 3022: Based on the environmental degradation label, the semantic segmentation map, and the bounding box information, select a matching enhancement algorithm and parameter configuration from the enhancement strategy library for each pixel of the original video frame to form an initial pixel-level enhancement scheme map.
[0139] 3023: Perform time-series filtering on the initial enhancement scheme map of the current original video frame and the enhancement scheme map of the previous original video frame to obtain a smoothed enhancement scheme map.
[0140] Steps 3022-3023 involve generating a pixel-level enhancement scheme diagram, which specifically corresponds to a process of multi-source fusion, layer-by-layer refinement, and dynamic decision-making. The following is a detailed explanation of its implementation steps: Step 1: Initialize the global basic enhancement scheme diagram The system first retrieves the corresponding global basic enhancement configuration from the policy library based on the environmental degradation label E_t output by the environmental classification network. Then, it creates a multidimensional data structure with the exact same size as the original image I_t, denoted as the scheme graph P_t_init. The parameter values from the retrieved global basic configuration are then filled into each pixel position of the scheme graph P_t_init.
[0141] For example, if the environment classifier determines that the current situation is "rainy night", the system queries the policy library and obtains the global configuration as "using a combination of Retinex algorithm A (intensity 0.6) and rain removal model B (intensity 0.5)". Then, in the initialized scheme graph P_t_init, the vector of each pixel is set to [algorithm ID=mixture, Retinex intensity=0.6, rain removal intensity=0.5,...].
[0142] Step 2: Perform regional adjustments based on the semantic segmentation map The system obtains the semantic segmentation map M_t output by the semantic segmentation network. Each pixel in this map is labeled with a semantic category (such as person, road, etc.). The system will traverse each pixel in the scheme map P_t_init, query the corresponding adjustment rule from the policy library according to its semantic category in M_t, and modify the parameter vector of that pixel.
[0143] For example, a "semantic adjustment matrix" can be created, with the number of rows equal to the number of semantic categories and the number of columns equal to the dimension of the parameter vector. For each pixel in the scheme diagram, based on its semantic label, the adjustment vector of the corresponding row is extracted from the adjustment matrix, and then a weighted sum or overwrite operation is performed with the current parameter vector of that pixel.
[0144] For example, continuing with the "rainy night" scene, for a pixel labeled "person," the policy library's rule is "Increase Retinex intensity by 15%, decrease rain removal intensity by 10%." The system will find the pixel's vector in P_t_init and modify it to [Retinex intensity = 0.6 * 1.15 = 0.69, rain removal intensity = 0.5 * 0.9 = 0.45, ...]. For a pixel labeled "sky," the rule might be "Disable Retinex, slight rain removal," and its vector might be modified to [Algorithm ID = rain removal only, Retinex intensity = 0, rain removal intensity = 0.3, ...]. After this step, the scheme graph P_t_semantic already contains different enhancement strategies based on content regions.
[0145] Step 3: Refined Coverage Based on Key Object Detection The system acquires bounding box information O_t output by the key object detection model. These bounding boxes define the regions that require the highest priority processing. The system will traverse each detected key object and refine the scheme map parameters within its covered pixel region.
[0146] Priority mechanism: Modifications in this step have the highest priority and will directly override the parameters generated in steps one and two. This ensures that even if a face is located in a road area, it will be treated as a face, not as road.
[0147] For each key object's bounding box, the system first determines the range of pixel coordinates it covers. Then, it queries the policy library for the dedicated enhancement scheme corresponding to that object type (e.g., "face"). The complete parameter vector of this dedicated scheme is then directly copied and applied to all pixels within the corresponding coordinate range in the scheme graph P_t_semantic.
[0148] For example, in a "rainy night" scene, a "face" region is detected. The dedicated scheme for "faces" in the policy library is "use face-specific network C, intensity 0.8, enable face bounding box color distortion protection". The system will locate all pixels within the face bounding box and uniformly set their vectors in the scheme graph, regardless of their previous values, to [Algorithm ID=face-specific network C, intensity=0.8, face bounding box color distortion protection=enabled, ...]. Similarly, if a "hand" is detected, the pixel vectors of its covered area will be overlaid with the "hand-specific scheme", which may include higher edge sharpening intensity to highlight any objects that may be held.
[0149] Step 4: Boundary Smoothing and Conflict Resolution After the layered processing described above, abrupt transitions may appear at the boundaries of different strategy regions in the scheme map. If such a scheme map is used directly for enhancement, the final image may exhibit obvious and unnatural block boundaries. Therefore, boundary smoothing is necessary. Although the priority mechanism resolves the main conflicts, more complex logic may still be required in some edge cases (such as one critical object being adjacent to another). For example, a "blending region" can be defined where the parameters of the two schemes are mixed in a certain proportion, or the blending ratio can be dynamically adjusted based on the distance of each pixel from the center of its respective object.
[0150] Step 5: Time series smoothing to maintain consistency Video is composed of consecutive frames. If the initial enhancement scheme map generated in step 3022 is used directly, the enhancement strategy between adjacent frames may change abruptly due to minor changes in the scene or accidental detection errors, resulting in flickering or discontinuous artifacts in the final output video. To address this issue, a temporal smoothing mechanism is introduced in this step. The system performs temporal filtering on the initial enhancement scheme map of the current frame and the enhancement scheme map of the previous frame (which has already undergone the same processing).
[0151] The system maintains a final enhancement scheme graph P_{t-1}_final for the previous frame. After generating the initial enhancement scheme graph P_t_smoothed for the current frame, the system performs temporal fusion of the two. A simple and effective method is exponential moving average: P_t_final = α * P_t_smoothed + (1 - α) * P_{t-1}_final Here, α is a smoothing factor between 0 and 1, controlling the influence of the current frame's enhancement scheme. A smaller α results in a more stable video but a slower response to scene changes; a larger α results in a faster response but may increase the risk of flickering. This α value can even be adaptively adjusted based on the movement speed of objects in the scene. After temporal filtering, the resulting smoothed enhancement scheme map is more continuous and stable in time, laying the foundation for subsequent generation of flicker-free enhanced videos.
[0152] After the above five steps, the system finally generates a highly refined, content-aware, and temporally coherent pixel-level enhancement scheme map P_t_final. This scheme map serves as the enhancement preprocessing instruction drawn for the current frame. It will be fed into a deep fusion network, which will precisely and pixel-level execute all the instructions on the map, ultimately outputting a high-quality, high-information-content enhanced image.
[0153] 303. Input the original video frame and the enhancement scheme map into a trained deep fusion network. Through the deep fusion network, using the enhancement scheme map as guiding information, perform adaptive and regional differentiation enhancement processing on the original video frame, and output a preliminary enhanced image. This invention employs an adaptive enhancement execution architecture based on a deep fusion network. Its core design concept is to transform the image enhancement process from an engineering problem of stacking algorithms into an intelligent problem of end-to-end learning. Instead of directly calling existing algorithm modules such as Retinex and DerainNet, it trains a unified and powerful deep neural network, enabling it to learn how to infer and draw a clear, high-quality enhanced image from the original degraded image based on instructions from a pixel-level enhancement scheme diagram.
[0154] This invention preferably employs a U-Net-like encoder-decoder structure, and makes targeted innovative modifications to it, enabling it to perfectly adapt to the needs of dual input streams (original image and scheme image) and pixel-level adaptive enhancement. The reasons for choosing the U-Net architecture are: firstly, its encoder-decoder structure is naturally suitable for image-to-image conversion tasks; secondly, its unique skip connections can directly pass the low-level, high-frequency detail features (such as edges and textures) extracted by the encoder to the decoder, which is crucial for preserving the true details of the image and avoiding excessive smoothing during the enhancement process. The detailed steps of step 303 are described below: 3031: Construct a U-Net-like deep fusion network, wherein the input layer of the deep fusion network is designed as a dual-channel network, used to receive the original video frame and the enhancement scheme diagram respectively.
[0155] U-Net is a convolutional neural network architecture that excels in image-to-image conversion tasks. Its encoder-decoder structure and skip connection design enable it to capture both global context and preserve local details when processing images. This invention modifies it into a dual-input architecture: one input channel receives the three-channel original RGB image (the original, video frame to be enhanced, I_t), and the other input channel receives the smoothed enhancement scheme map (P_t_final) generated in step 3023.
[0156] 3032: In the encoder part of the deep fusion network, content features and policy features are extracted from the original video frame and the enhancement scheme graph, respectively.
[0157] The encoder consists of a series of convolutional and downsampling layers, processing the two inputs separately. For the raw video frame, the encoder extracts content features such as edges, textures, and the shape and structure of objects. For the enhancement scheme map, the encoder extracts policy features, i.e., the pattern and spatial distribution of enhancement instructions. For example, policy features might encode instructions such as "apply strong dehazing in this area and weak sharpening in that area."
[0158] (1) Extraction of content features from the original image The raw video frame I_t first enters the network's content encoder. This encoder is typically composed of a series of convolutional layers, activation layers (such as ReLU), and downsampling layers (such as max pooling or strided convolution). Its main task is to deeply understand the visual content of the image.
[0159] Shallow networks: In the initial stages of the encoder, the convolutional kernels are small and used to capture basic visual elements such as edges, corners, color gradients, and textures. These are the foundation of an image.
[0160] Deep networks: As network depth increases, the size of the feature map gradually decreases through downsampling operations, but the number of channels (i.e., feature dimension) gradually increases. Deep networks can learn more abstract and higher-level semantic information, such as the outline of an object, the combination of parts, and even the category of the object.
[0161] After complete processing by the content encoder, the original image I_t is transformed into a set of multi-scale, hierarchical content feature maps. This set of feature maps encodes comprehensively the information of "what" and "what" the original image "has" in a compact form.
[0162] (2) Enhanced strategy feature extraction of solution graph Meanwhile, the enhanced scheme graph P_t_final is fed into the network's policy encoder. This encoder has a similar structure to the content encoder, but its function and objectives are completely different.
[0163] Instruction Decoding: Each pixel location in the enhancement scheme graph P_t_final contains a high-dimensional parameter vector, such as [Algorithm ID, Enhancement Strength_1, Enhancement Strength_2, Special Switch_1, ...]. When processing this data, the convolutional kernel of the policy encoder learns not visual patterns, but the spatial relationships and patterns between these parameters. For example, it can understand that a continuous region with the same algorithm ID means that a large object needs to be processed uniformly.
[0164] Policy abstraction: Through layer-by-layer convolution and downsampling, the policy encoder abstracts specific, discrete parameter instructions into higher-level, more instructive policy feature maps. For example, it may merge two specific instructions, "use high-intensity low-light enhancement for the face region" and "use high sharpening for the hand region," into an abstract policy feature that "focuses on key details of the person."
[0165] 3033: In the bottleneck layer of the deep fusion network, the content features and policy features are deeply fused.
[0166] The bottleneck layer, the deepest part of the network with the smallest feature map, is the key node for information fusion. Here, the system fuses features extracted from two input paths. The fusion methods can be varied; for example, the two features can be concatenated along the channel dimension and then mixed through a convolutional layer; or an attention mechanism can be used, where policy features are used as weights to modulate the importance of content features. The goal of this step is to teach the network how to adjust and enhance content according to the policy, achieving intelligent image processing.
[0167] Once content features and policy features are extracted from their respective encoding paths, the core of the entire network—the deep fusion mechanism—begins to function. This is crucial for achieving adaptation, aiming to perfectly combine "what to do" (policy) with "what to do with" (content). This invention can employ various advanced fusion mechanisms: (1) Feature-level spatial adaptive modulation The core idea of adaptive modulation is to allow policy features to modulate or calibrate content features, so that the content features dynamically change their expression according to policy instructions at different spatial locations.
[0168] At the bottleneck layer of the network (i.e., the deepest layer of the encoder where the feature map size is smallest), the system aligns the content feature map and the policy feature map. The policy feature map is then used to generate a set of spatially varying gating signals. The gating signal can be a tensor with the same number of channels and spatial size as the content feature map, with values between 0 and 1. The content features are then multiplied element-wise with this gating signal. In regions where the policy map indicates a strong need for enhancement, the gating signal value is close to 1, and the content features are almost completely preserved or even amplified; while in regions where the policy map indicates no need for processing, the gating signal value is close to 0, and the corresponding content features are significantly suppressed. In this way, the network learns to "follow instructions," processing only the image information in the regions of interest.
[0169] (2) Attention-based guided fusion A spatial attention module is constructed using the policy feature map as the "query" and the content feature map as the "key" and "value." This module calculates the "relevance score" between each location on the content feature map and the policy instruction. This relevance score map is essentially an attention map, highlighting regions where the image content is highly relevant to the enhancement policy. For example, if the policy map emphasizes "hands," the attention map will produce a high response at the location of the hand in the original image. Finally, this attention map is used to weight the content feature map, allowing the network to allocate more "effort" and "resources" to the regions the policy focuses on during decoding and reconstruction.
[0170] 3034: In the decoder part of the deep fusion network, the fused features are upsampled to generate the preliminary enhanced image.
[0171] After deep fusion, a highly condensed fusion feature map containing both "content" and "strategy" is obtained. The next task is to use a decoder to "restore" or "redraw" these abstract feature maps step by step into a clear, enhanced image.
[0172] The decoder consists of a series of upsampling and convolutional layers. It restores the resolution of the original image by progressively upsampling the deep feature maps fused from the bottleneck layer. U-Net's skip connections directly pass shallow content features from different levels in the encoder to the corresponding upsampling modules in the decoder. This supplements the deep, semantically rich features with shallow, high-resolution details (such as fine edges), helping the decoder reconstruct a detailed, artifact-free enhanced image. Ultimately, the decoder outputs a preliminary enhanced image. This image has undergone high-quality adaptive enhancement according to the intelligently defined scheme, but minor inconsistencies between frames may still exist.
[0173] The fused feature maps received by the decoder have already been imprinted with enhancement instructions. Therefore, the decoder's behavior when reconstructing each pixel is no longer fixed but adaptive. During training, the decoder learns to map different fused feature patterns to specific pixel operations. For example, when the decoder processes a fused feature in a certain region, and its policy strongly indicates "low-light enhancement," the corresponding neurons inside the decoder are activated, performing operations similar to "increasing brightness and adjusting contrast." When the policy indicates "removing rain," another set of neurons is activated, performing operations similar to "removing rain streaks and restoring the background."
[0174] Since the fusion mechanism (such as spatial adaptive modulation) is inherently smooth, the features passed to the decoder also change gradually between different policy regions. Therefore, when reconstructing boundaries, the decoder can naturally and smoothly transition from one processing mode (such as face enhancement) to another mode (such as background deraining), thus avoiding the harsh boundaries and artifacts common in traditional methods.
[0175] The following is an explanation of the training process for deep fusion networks: (1) Construction of training dataset Training this network requires a special dataset, where each sample contains a triplet: {original degraded image I_t, pixel-level augmentation scheme image P'_t, target high-quality image I_gt}.
[0176] I_t: Can be obtained by taking pictures under various weather and lighting conditions, or by algorithmically degrading sharp images. P'_t: Can be automatically generated based on I_t and I_gt using a teacher model (i.e., the previously designed policy library rules). I_gt: Requires truly high-quality sharp images that need to be collected manually or taken under ideal conditions, serving as the ultimate learning goal.
[0177] (2) Design of loss function To guide the network in learning correct reinforcement behaviors, a composite loss function needs to be designed, which may include one or more of the following losses: 1) Reconstruction loss: Calculate the pixel-level difference (such as L1 or L2 loss) between the enhanced image I'_t output by the network and the target high-quality image I_gt.
[0178] 2) Perceptual loss: In order to make the generated images more visually realistic and more in line with human visual perception, a pre-trained image classification network (such as VGGNet) can be introduced to calculate the difference between I'_t and I_gt in the deep features of the network.
[0179] 3) Adversarial Loss: A discriminator network can be introduced to form a Generative Adversarial Network (GAN) framework. The generator (i.e., a deep fusion network) strives to generate realistic images, while the discriminator tries to distinguish between real, high-quality images and generated images. This game-theoretic approach can greatly improve the clarity and realism of the generated images.
[0180] 4) Instruction Consistency Loss: Its function is to ensure that the network actually follows the instructions of the scheme diagram P'_t. Specifically, this can be achieved by inputting the scheme diagram P'_t into an auxiliary network, predicting its expected effect, and comparing it with the actual output of the main network. Alternatively, it can be achieved by directly analyzing the characteristics of the output image I'_t in different regions (such as brightness, contrast, and noise level) to see if it matches the corresponding instruction parameters in the scheme diagram P'_t.
[0181] 304. Perform temporal fusion processing on the preliminary enhanced image to output temporally coherent enhanced video frames.
[0182] If video streams are enhanced frame by frame independently, severe visual artifacts such as time flickering are almost inevitable. To break the isolated mode of frame-by-frame processing, this embodiment introduces information from the timeline into the enhancement decision of the current frame. The system no longer generates the final result solely based on the information of the current frame, but looks back and uses information from the previous frame or even several previous frames to predict and calibrate the processing result of the current frame, thereby achieving a smooth and stable output. This embodiment proposes a dual fusion mechanism, which performs temporal optimization at the "instruction level" (pixel-level enhancement scheme map) and the "result level" (enhanced image), forming a three-dimensional, multi-layered anti-flicker system. The first fusion refers to the temporal smoothing of the pixel-level enhancement scheme map. This is the first layer of the dual fusion mechanism of this invention, which performs preprocessing at the instruction level to prevent abrupt changes in the enhancement strategy from the source. Specifically, the initial pixel-level enhancement scheme map P_t generated in the current frame without temporal processing, and the final scheme map P'_{t-1} of the previous frame after temporal smoothing, are subjected to temporal filtering, as detailed in the embodiment description of step 3023.
[0183] The detailed implementation of step 304 is described below: 3041: Calculate the optical flow field between the current original video frame and the previous original video frame.
[0184] This invention preferably employs optical flow for motion estimation. The optical flow field is a two-dimensional vector field that describes the direction and velocity of motion of each pixel in a video sequence from the previous frame to the current frame. Calculating optical flow is a classic problem in computer vision and can be accurately estimated using deep learning-based methods such as FlowNet2 and PWC-Net. The optical flow field accurately captures the motion information of the image.
[0185] The system simultaneously inputs the current original frame I_t and the previous original frame I_{t-1} into a trained optical flow computation network (such as PWC-Net, RAFT, etc.). This network, through complex convolutional operations and cost volume aggregation, efficiently outputs an optical flow field F_t with the same size as the original image and two channels. In this optical flow field, the two-dimensional vector (u, v) at each position (i, j) represents the pixel at position (i, j) in the original frame I_{t-1}, which has moved to a new position (i+u, j+v) in the current frame I_t. This optical flow field F_t serves as the guiding information for all subsequent temporal operations, precisely establishing the geometric bridge between the two frames.
[0186] 3042: Using the optical flow field, the enhanced video frame output from the previous frame is warped and aligned to the viewpoint of the current frame.
[0187] After obtaining the optical flow field, the system performs a geometric transformation on the previously processed "enhanced video frame." Specifically, based on the motion vector of each pixel indicated by the optical flow field, the system "moves" or "distorts" the pixels from the previous frame to their new positions in the current frame. After this step, the enhanced image from the previous frame is "aligned" to the geometric viewpoint of the current frame. In this way, static objects in the two frames perfectly overlap, with only newly appearing or occluded areas showing differences.
[0188] This step operates on the enhanced image I'_{t-1} of the previous frame. The final result after complete temporal optimization is used here because it contains the most stable historical information. The system uses the optical flow field F_t to reverse-distort the enhanced image I'_{t-1} of the previous frame. This process can be visualized as follows: according to the direction of the optical flow field F_t, each pixel in the enhanced image I'_{t-1} of the previous frame is "moved" or "sampled" back to its corresponding position in the current frame I_t along the opposite direction of its motion vector. After this operation, a new image is obtained, namely the aligned historical enhanced image I'_aligned. The content of this image is geometrically perfectly aligned with the current frame I_t, but its visual style and quality inherit the stability of the previous frame. For example, if a person moves from frame t-1 to frame t, then in the aligned image I'_aligned, the person's image will appear precisely at their position in frame t, but their appearance (brightness, sharpness, etc.) will still be the stable state after the enhancement of frame t-1.
[0189] 3043: The preliminary enhanced image is weighted and fused with the aligned previous enhanced video frame to output a temporally coherent enhanced video frame.
[0190] The second layer of fusion is temporal fusion to enhance the results. This is the second layer of the dual fusion mechanism, used to perform fusion at the final "result" level to ensure the smoothness of the output image. Specifically, the deep fusion network generates the initial enhanced image I'_t of the current frame based on the smoothed scheme map P'_t, and the aligned historical enhanced image I'_aligned. The system performs a weighted fusion of these two images to generate the final enhanced video frame I''_t. The fusion formula is as follows: I''_t = α * I'_t + (1 - α) * I'_aligned Here, α is the fusion weight factor. The weight α can be dynamically adjusted based on various factors, such as motion confidence, content changes, and enhancement intensity.
[0191] This further fusion step is equivalent to using the stable result of the previous frame to "constrain" and "smooth" the enhanced result of the current frame. Even if the deep fusion network outputs an I'_t that is slightly different in style from the previous frame for some reason, this difference will be effectively smoothed out through fusion with I'_aligned, and the final output I''_t will exhibit a highly stable and temporally coherent transition in time.
[0192] After completing all the above steps, the system obtains the final high-quality, temporally coherent enhanced video frame I''_t. I''_t is then sequentially fed into the video stream encoder to form the final high-quality video stream for use by the downstream hazard warning system. To prepare for the processing of the next frame, the system needs to update the temporal buffer. Specifically, this includes saving the smoothed scheme map P'_t of the current frame as P'_{t-1} for the next frame, and saving the final enhanced image I''_t of the current frame as I''_{t-1} for the next frame. This forms a complete, continuously running closed-loop processing flow.
[0193] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the public place danger warning model training method described in the above embodiments, or to perform the steps of the public place danger warning method described in the above embodiments.
Claims
1. A training method for a public place hazard early warning model, characterized in that, Includes the following steps: Acquire historical monitoring data containing confirmed hazardous events and synchronized audio; Multi-target cross-camera trajectory tracking is performed on the historical monitoring data, and audio and video features are extracted at fixed time intervals to form regional monitoring vectors. Then, the regional monitoring vectors of multiple consecutive time points are arranged in chronological order to form time-series monitoring vectors. The occurrence time of each dangerous event in the historical monitoring data is marked, and an analysis interval with a preset duration is defined that is located before the occurrence time. If the end time of the time-series monitoring vector is located within the analysis interval, it is marked as the first sample; otherwise, it is marked as the second sample. Based on the time-series monitoring vector and the corresponding sample identifier, the pre-trained model is trained to obtain a public place danger early warning model for outputting the probability of future dangerous events. The step of performing multi-target cross-camera trajectory tracking on the historical monitoring data, extracting audio and video features at fixed time intervals to form regional monitoring vectors, and then arranging the regional monitoring vectors from multiple consecutive time points in chronological order to form a time-series monitoring vector includes: Initialize a multi-object tracking engine and load a trained object detection model and a trained object re-identification model; For each frame of video image in the historical monitoring data, the target detection model is run to detect all moving targets and obtain the bounding box of each moving target; Using the target re-identification model, a depth appearance feature vector is calculated for each detected moving target, and combined with motion information, a unique identifier is assigned to the same moving target appearing in consecutive frames for generating single-camera trajectories; By applying cross-camera re-identification technology, the apparent feature vectors of all moving targets appearing in the fields of view of different cameras are compared, and the trajectories of single cameras belonging to the same moving target are spatiotemporally correlated to construct the complete cross-regional spatiotemporal trajectory of the moving target within the global monitoring area. At fixed time intervals, multimodal feature snapshots are taken of each monitoring area at the current moment. Specifically, this includes: traversing all tracked moving targets within the monitoring area; for each moving target, extracting video features from the corresponding image area using a trained attribute recognition network; and extracting audio features from the synchronized audio signal within the current time window using a trained audio classification network. All extracted video and audio features are vectorized, and the video feature vectors of all moving targets in the same monitoring area at the same time and the global audio feature vector are concatenated to form a high-dimensional regional monitoring vector. The regional monitoring vectors of N consecutive time points are arranged in chronological order to form a time-series monitoring vector of length N, which is then used as a basic training sample for the pre-trained model.
2. The training method for a public place hazard early warning model according to claim 1, characterized in that, The process of labeling the occurrence time of each dangerous event in the historical monitoring data and defining an analysis interval that is prior to the occurrence time and has a preset duration, wherein if the end time of the time-series monitoring vector is within the analysis interval, it is marked as a first sample, otherwise it is marked as a second sample, includes: Deploy a pre-trained hazard event recognition model, wherein the hazard event recognition model is a multimodal classifier based on a convolutional neural network-long short-term memory network; Using the aforementioned hazardous event identification model, the historical monitoring data is scanned and analyzed frame by frame, and the specific time point, specific monitoring area, and event category of each type of hazardous event are marked. For each marked hazardous event, a hazardous event analysis interval is defined that is located before the time point of occurrence corresponding to the hazardous event. The analysis interval includes the start time, end time and duration. Create an empty training dataset and iterate through all time-series monitoring vector samples; For each time series monitoring vector sample, check whether the time series monitoring vector sample falls within any defined analysis interval at the end time; If the time-series monitoring vector sample falls within the analysis interval of the dangerous event at the end time, then an M-dimensional vector first sample label is created for the time-series monitoring vector sample, where M is the total number of dangerous event categories, and the component corresponding to the dangerous event in the first sample label is set to 1, and the other components are set to 0. If the time-series monitoring vector sample does not fall within the analysis interval of any dangerous event at the end time, then a second sample label is created for the time-series monitoring vector sample, which is an M-dimensional vector with all components being 0. All the constructed time-series monitoring vector samples are paired with their corresponding sample identifiers according to time points to form a complete model training dataset.
3. The training method for a public place hazard early warning model according to claim 1 or 2, characterized in that, The step of training the pre-trained model based on the time-series monitoring vector and the corresponding sample identifier to obtain a public place hazard early warning model for outputting the probability of future dangerous events includes: A long short-term memory network is selected as the basic architecture of the pre-trained model, and the network structure parameters are configured. A fully connected layer is connected after the output layer of the long short-term memory network, and a preset activation function is used to make the pre-trained model output a multi-dimensional probability vector, where the value of each component represents the probability of the corresponding category of dangerous event occurring in the future. The prepared training dataset is divided into a training set and a validation set according to a preset ratio; Set the binary cross-entropy as the loss function of the model, select the preset optimizer as the optimization algorithm, and set the initial learning rate and weight decay coefficient. The training process is initiated by inputting the training set data into the pre-trained model in batches, performing forward propagation calculations to calculate the loss between the predicted probability and the true probability, then calculating the gradient through the backpropagation algorithm, and updating the network weights using the preset optimizer. After each training cycle, the model's performance is evaluated using the validation set, and the loss value and accuracy on the validation set are monitored. If the loss value on the validation set no longer decreases over multiple consecutive training cycles, then the model training is terminated, the optimal model parameters are saved, and the trained public place danger early warning model is obtained.
4. The training method for a public place hazard early warning model according to claim 1, characterized in that, The step of taking multimodal feature snapshots of each monitored area at a fixed time interval includes: When performing video feature extraction, at fixed time intervals, for each tracked moving target, the trained attribute recognition network is used to extract basic attributes and behavioral state attributes from the image region corresponding to the moving target. When performing audio feature extraction, the event sound features and environmental sound features are extracted using a trained audio classification network at fixed time intervals for synchronous audio signals. The extracted categorical attributes are vectorized using one-hot encoding, the extracted numerical attributes are normalized and vectorized, and all vectorized video and audio features are concatenated to form a high-dimensional feature vector. The types of moving targets include personnel targets and vehicle targets. The basic attributes of personnel targets include: gender, age group, height, and body type. The behavioral status attributes of personnel targets include: whether they are carrying luggage, whether they are carrying a suspected knife, whether they are masked, the number of people in their group, whether they are running, whether they are exhibiting abnormal limb movements, and whether they have fallen. The basic attributes of vehicle targets include: license plate number and vehicle type. The behavioral status attributes of vehicle targets include: whether they have entered a restricted area and whether they have turned on their hazard lights. The event sound features include whether screams, cries for help, loud arguments, the sound of breaking glass, and vehicle horns are detected. The environmental sound features include the background volume decibel value and the noise level of the crowd.
5. A method for early warning of danger in public places, characterized in that, Includes the following steps: Video streams and synchronous audio streams are acquired in real time from surveillance cameras in various monitoring areas of public places, and the video streams are preprocessed with image enhancement. Multi-target cross-camera trajectory tracking is performed on the video stream after image enhancement preprocessing, and video features of the video stream and audio features of the audio stream are extracted at fixed time intervals; The extracted video and audio features are concatenated to obtain the area monitoring vector at the current moment. The latest multiple area monitoring vectors are then arranged in chronological order to form a time-series monitoring vector. The time-series monitoring vector is input into a pre-trained public place danger early warning model, which outputs probability vectors of various dangerous events that will occur in each monitoring area of the public place in the future. Determine whether there is a probability component in the probability vector that is greater than the warning threshold; If there is a probability component greater than the warning threshold, a danger warning is triggered, generating warning information containing the type of dangerous event, the area where the danger occurred, and a summary of key features, and security joint control is initiated. The image enhancement preprocessing of the video stream includes: Each original video frame of the video stream is sequentially subjected to environmental degradation classification, semantic content segmentation, and key object detection to obtain scene cognition and analysis results; Based on the scene recognition and analysis results, a specified enhancement algorithm is generated for each pixel of the original video frame, and a corresponding enhancement scheme diagram is configured. The original video frame and the enhancement scheme map are input into a trained deep fusion network. The deep fusion network, guided by the enhancement scheme map, performs adaptive and regional differentiation enhancement processing on the original video frame and outputs a preliminary enhanced image. The initially enhanced image is subjected to temporal fusion processing to output temporally coherent enhanced video frames.
6. The public place danger early warning method according to claim 5, characterized in that, The process of sequentially performing environmental degradation classification, semantic content segmentation, and key object detection on each original video frame of the video stream to obtain scene cognition and analysis results includes: Each original video frame of the video stream is sequentially input into a trained environment classification network, which outputs an environment degradation label representing the main degradation type of the current scene; each original video frame of the video stream is sequentially input into a trained real-time semantic segmentation network, which generates a semantic segmentation map that identifies different semantic regions; each original video frame of the video stream is sequentially input into a trained key object detection model, which outputs a set of bounding box information containing the location and category of key objects. The step of generating a specified enhancement algorithm for each pixel of the original video frame based on the scene recognition and analysis results, and configuring a corresponding enhancement scheme diagram, includes: An enhancement strategy library containing various basic enhancement algorithms and their different parameter configurations is pre-established. Based on the environmental degradation label, the semantic segmentation map, and the bounding box information, a matching enhancement algorithm and parameter configuration are selected from the enhancement strategy library for each pixel of the original video frame to form an initial pixel-level enhancement scheme map. The initial enhancement scheme map of the current original video frame and the enhancement scheme map of the previous original video frame are subjected to temporal filtering to obtain the final enhancement scheme map.
7. The public place danger early warning method according to claim 5, characterized in that, The process involves inputting the original video frame and the enhancement scheme map into a trained deep fusion network. Using the enhancement scheme map as guiding information, the deep fusion network performs adaptive and region-differentiated enhancement processing on the original video frame, outputting a preliminary enhanced image. A U-Net-like deep fusion network is constructed. The input layer of the deep fusion network is designed with dual channels to receive the original video frame and the enhancement scheme map, respectively. In the encoder part of the deep fusion network, content features and policy features are extracted from the original video frame and the enhancement scheme map, respectively. In the bottleneck layer of the deep fusion network, the content features and the policy features are deeply fused. In the decoder part of the deep fusion network, the fused features are upsampled to generate the preliminary enhanced image. The step of performing temporal fusion processing on the preliminary enhanced image to output temporally coherent enhanced video frames includes: calculating the optical flow field between the current original video frame and the previous original video frame; using the optical flow field, warping and aligning the enhanced video frame output from the previous frame to the viewpoint of the current frame; and performing weighted fusion of the preliminary enhanced image and the aligned previous enhanced video frame to output temporally coherent enhanced video frames.
8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the public place danger early warning model training method as described in any one of claims 1-4, or executes the public place danger early warning method as described in any one of claims 5-7.
Citation Information
Patent Citations
Safety risk analysis method and system based on intelligent camera
CN119478755A
Monitoring image processing method and system for high-precision target tracking
CN120676252A