Community population infection risk-oriented multi-source behavior data intelligent early warning method

By dividing real-time video streams into time windows and tracking human targets, and using a historical tag ID priority strategy, high-frequency risk behaviors can be identified and warned. This solves the problem of the difficulty in quickly identifying high-frequency risk individuals in densely populated places with elderly people, and enables earlier risk warnings and more efficient management.

CN121789145APending Publication Date: 2026-04-03XUZHOU MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly identify and provide timely warnings of high-frequency risk behaviors in community settings frequented by the elderly, especially when high-risk behaviors are recurring and prolonged among the elderly population. Existing systems lack prioritization strategies for individuals who have previously exhibited risk signs, resulting in insufficient analysis of key targets.

Method used

By acquiring real-time video streams and dividing them into time windows, human target detection and multi-target tracking are performed to obtain a unique tracking ID. Using the IDs already marked within the historical time window as priorities and combining spatial distance sorting, behavioral characteristics are analyzed and risk scores are calculated to achieve the identification and early warning of high-risk individuals.

Benefits of technology

It has achieved stable identification of high-frequency risk individuals and reduced false alarms, improving real-time performance, accuracy and interpretability, and enabling earlier detection of potential cluster transmission risks and timely warnings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789145A_ABST
    Figure CN121789145A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source behavior data intelligent early warning method for community crowd infection risks. The method comprises the following steps: acquiring a real-time video stream; performing human body target detection and multi-target tracking on each frame of image to obtain a unique tracking ID; calling the unique tracking ID marked in the first historical time window, judging whether the unique tracking ID appears in the current time window or not, if so, taking the unique tracking ID as a first priority, and sorting the rest of the unique tracking ID in an ascending order according to the spatial distance from the unique tracking ID to the first priority; and performing region-of-interest division according to a priority sequence and judging whether a marking event feature is met or not, if so, marking the unique tracking ID, counting the marking times of the unique tracking ID in a second historical time window, screening out all unique tracking IDs exceeding a safety threshold, performing comprehensive calculation to obtain a risk score, and performing early warning according to the risk score. According to the method, ROI and action sequence analysis is preferentially performed on the high-risk associated object, so that the computing power consumption caused by equivalent analysis of all personnel is reduced, and the real-time early warning capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of behavioral risk detection technology, specifically to an intelligent early warning method based on multi-source behavioral data for community population infection risk. Background Technology

[0002] In the context of community-based elderly care and the "home-community" integrated elderly care service model, public activity venues such as community chess and card rooms, dance studios, calligraphy and painting studios, and senior activity centers have become important spaces for the daily social and cultural activities of the elderly. These venues typically feature long periods of time spent in close proximity, frequent face-to-face interactions, and significant fluctuations in activity intensity and respiratory behavior. If a carrier of a respiratory infectious disease (such as influenza or coronavirus) is present, there is a high risk of rapid cluster transmission. Meanwhile, authoritative public health institutions generally point out that the elderly (e.g., ≥65 years old) are among the high-risk groups for severe illness or adverse outcomes after respiratory viral infections. Therefore, early identification and intervention of infection risks in community venues frequented by the elderly have greater public health significance and management value.

[0003] From a management practice perspective, community risk control in senior activity rooms often faces the following practical challenges: On the one hand, relying solely on manual inspections, verbal reminders, or post-incident investigations results in issues such as discontinuous coverage, delayed response, difficulty in quantification, and difficulty in tracing high-risk individuals and key behavioral segments; on the other hand, senior activities have their own characteristics, such as: some seniors may have hearing or reaction time decline, differing understandings of protective measures, improper mask wearing, and prolonged close-range conversations / sitting in enclosed spaces, making risky behaviors more likely to recur and last longer, thus placing higher demands on real-time monitoring systems: not only must they determine whether the risk in the venue has increased, but they must also be able to identify which individuals repeatedly trigger risky behaviors and whether the risk is accumulating.

[0004] While existing technologies include intelligent analysis approaches based on video surveillance, most are geared towards open communities. These technologies identify infection pathways by tracking the movement of target individuals. While these technologies can achieve continuous monitoring without direct contact with the elderly and are well-suited for community activity rooms such as chess and card rooms and dance studios, they still have shortcomings for applications in community activity rooms primarily frequented by the elderly.

[0005] Because of the relatively enclosed environment, existing systems mostly focus on location-level density / event statistics. Senior activity rooms often experience peak periods of activity (e.g., fixed times for chess, card games, or dancing), and some events involve high-risk behaviors repeatedly exhibited by the same senior individual over a longer timescale. If the system performs equal analysis on all individuals (frame-by-frame / window-by-window traversal of ROI and action sequences), it may lead to difficulties in quickly identifying and focusing on high-frequency risk individuals. Furthermore, the lack of a prioritization strategy for those who have previously shown signs of risk prevents the system from focusing on key individuals and issuing alerts first, resulting in insufficient analysis of critical targets. Summary of the Invention

[0006] To address the problem that existing technologies struggle to quickly identify and focus on individuals at high risk, this invention provides an intelligent early warning method based on multi-source behavioral data for community population infection risk.

[0007] To achieve the above objectives, the technical solution of the present invention is as follows:

[0008] Firstly, this application discloses an intelligent early warning method based on multi-source behavioral data for community population infection risk, comprising the following steps:

[0009] Acquire real-time video streams captured by cameras within the target public activity area and divide the video streams into time windows;

[0010] Perform human target detection and multi-target tracking on each frame of the video stream within the current time window to obtain a unique tracking ID for the human target;

[0011] Retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, take it as the first priority, and sort the rest in ascending order of spatial distance from the first priority;

[0012] The human target is divided into regions of interest in order of priority and the action sequence of the region of interest is constructed. The action sequence is then subjected to feature extraction to determine whether it meets the characteristics of the labeled event. If it does, the unique tracking ID of the human target is labeled and updated to the database.

[0013] Based on the unique tracking ID marked in the current time window, the number of times it was marked in the second historical time window is counted from the updated database and sorted in descending order. All unique tracking IDs that exceed the safety threshold are filtered out, and their behavioral characteristics are analyzed one by one in descending order. The risk score is calculated and matched with the corresponding risk level. The corresponding image fragment is then sent to the monitoring terminal for early warning.

[0014] The second historical time window is larger than the first historical time window.

[0015] Secondly, this application discloses a multi-source behavioral data intelligent early warning system for community population infection risk, including a data acquisition module, a target tracking module, a priority setting module, an event judgment module, and a data output module.

[0016] The data acquisition module is used to acquire real-time video streams captured by cameras in the target public activity area and to divide the video streams into time windows;

[0017] The target tracking module is used to perform human target detection and multi-target tracking on each frame of the video stream within the current time window, and obtain a unique tracking ID for the human target;

[0018] The priority setting module is used to retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, it is set as the first priority, and the rest are sorted in ascending order according to the spatial distance from the first priority.

[0019] The event judgment module is used to divide the human target into regions of interest in order of priority and construct the action sequence of the regions of interest. The action sequence is then used to extract features and determine whether it meets the characteristics of the event to be marked. If it does, the unique tracking ID of the human target is marked and updated to the database.

[0020] The data output module is used to count the number of times a unique tracking ID is marked in the second historical time window from the updated database based on the unique tracking ID marked in the current time window, sort it in descending order, filter out all unique tracking IDs that exceed the safety threshold, analyze their behavioral characteristics one by one in descending order, calculate the risk score, match the corresponding risk level, and send the warning to the monitoring terminal in combination with the corresponding image fragment.

[0021] The second historical time window is larger than the first historical time window.

[0022] Thirdly, this application discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned intelligent early warning method for multi-source behavioral data on infection risk in community populations.

[0023] Fourthly, this application discloses an intelligent early warning device based on multi-source behavioral data for community population infection risk, including at least one camera, a memory, and a processor.

[0024] The cameras are installed in the target public activity area to collect real-time video streams of the target public activity area;

[0025] The memory is used to store computer programs and a pre-set database, which at least stores unique tracking IDs and their tagging information within the first and second historical time windows.

[0026] The processor communicates with the camera and the memory. The processor is used to call and execute the computer program stored in the memory. When the program is executed by the processor, it implements the steps of the aforementioned intelligent early warning method for multi-source behavioral data on the risk of infection in the community population.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] 1. This application obtains a unique tracking ID through human detection and multi-target tracking, enabling the behavior of the same person to be continuously correlated in different frames and time windows. This provides a basis for subsequent historical tag retrieval, frequency statistics, and risk assessment, thereby avoiding only coarse-grained judgment at the location level. The application uses the tagged IDs in the first historical time window as the first priority and sorts other targets according to their spatial distance, prioritizing ROI and action sequence analysis of high-risk correlated objects to meet the need to focus on key targets first.

[0029] 2. This application uses the second historical time window to count the number of times an event is marked and compares it with a safety threshold to screen out objects that frequently trigger risk events. Then, it calculates a risk score based on behavioral characteristics and matches the risk level, thus upgrading the process from event detection to risk assessment and graded handling. By prioritizing historically high-risk individuals and sorting them by spatial distance, the system pays more attention to people who may have a higher transmission association and their neighboring population, which is conducive to discovering potential cluster transmission risks earlier and providing timely warnings. Attached Figure Description

[0030] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:

[0031] Figure 1 This is a flowchart illustrating the intelligent early warning method based on multi-source behavioral data for community population infection risk, as described in this invention.

[0032] Figure 2 Based on Figure 1 The flowchart for judging the marked event;

[0033] Figure 3 Based on Figure 1 The flowchart for calculating the risk score;

[0034] Figure 4 Based on Figure 3 A flowchart for fusing all unique tracking IDs with corrected single risk scores;

[0035] Figure 5 Based on Figure 3A flowchart for calculating the risk index of masks within the contact area;

[0036] Figure 6 Based on Figure 3 A flowchart for risk identification based on environmental factors;

[0037] Figure 7 Based on Figure 1 YOLO model workflow diagram;

[0038] Figure 8 Based on Figure 1 A schematic diagram illustrating the target location relationships in the application scenario;

[0039] Figure 9 This is a block diagram of a multi-source behavioral data-based intelligent early warning system for community population infection risk, as described in Example 2.

[0040] Figure 10 This is a schematic diagram illustrating the collaborative operation of a multi-source behavioral data-based intelligent early warning device for community infection risk, as described in Example 4. Detailed Implementation

[0041] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.

[0042] In existing technologies, infection risk identification in public activity venues is mainly based on single time segments or venue-level statistics: the video side often only performs density, contact or action identification on the current frame, and the results are usually output in the form of risk values, but lack consistent binding with individual identity and linkage with historical data; at the same time, there is a lack of closed-loop coordination between event data, trajectory data and historical records, which makes it difficult to accumulate and trace the risk behaviors of the same person in different time windows, and thus it is impossible to distinguish between occasional false detections and continuous high risks from the data level.

[0043] To address the aforementioned issues, this application constructs an innovative closed loop using a collaborative link of "time window data—tracking ID data—historical marker database—spatial relationship data—action sequence feature data—statistical and hierarchical data". First, the real-time video stream is organized into computable units according to time windows, and a unique tracking ID is generated for each frame of human detection and multi-target tracking within the window, so that behavioral data can be stably attributed to specific individuals. Then, the database retrieves the marked IDs in the first historical time window and matches them with the IDs in the current window. Historical risk records are directly used for priority scheduling in the current analysis, and computational resources are focused on the targets most likely to generate propagation associations by combining spatial distance sorting. Subsequently, action sequences are constructed for the ROI of priority targets and features are extracted. Once a marking event is met, it is written back to the database with a unique ID, forming a closed-loop data flow of "detection-marking-update". Further, the number of markings is counted and thresholded in the second historical time window (greater than the first historical time window). Short-term recurring clues are combined with long-term accumulated evidence to achieve stable identification of high-frequency risk individuals and reduce false alarm gating. Finally, for the targets that pass the screening, a risk score is comprehensively calculated, a risk level is matched, and a warning is pushed with the corresponding image fragment, thereby improving real-time performance, accuracy, interpretability, and handling efficiency at the data collaboration level.

[0044] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0045] Example 1

[0046] like Figure 1 The diagram illustrates an intelligent early warning method based on multi-source behavioral data for assessing infection risk in community populations, comprising the following steps:

[0047] S100. Acquire real-time video streams captured by cameras within the target public activity area and divide the video stream into time windows.

[0048] Specifically:

[0049] Select at least one fixed camera (which can be an IPC network camera / PTZ camera / wide-angle camera) within the target public activity area, complete the installation location calibration and data acquisition parameter configuration. The data acquisition parameters should include at least: video resolution, frame rate (FPS), bit rate, encoding format (such as H.264 / H.265), video stream protocol (such as RTSP / HTTP-FLV / WebRTC), and timestamp acquisition method (camera-side timestamp or unified time synchronization at the access end). Establish a mapping relationship between camera ID and venue area (such as activity room, chess and card room, etc.) for subsequent multi-camera data management and alarm location.

[0050] The real-time video stream is accessed through the video stream pull module, the video stream is decoded into a continuous frame sequence, and frame metadata is added to each frame. The frame metadata includes at least: frame timestamp t, camera ID, frame sequence number, and image width and height information; at the same time, a circular buffer queue or message queue is established to cache video frames in the most recent period to meet the frame retrieval requirements of subsequent image segment return warnings.

[0051] Time synchronization correction for real-time video streams: When timestamp jumps, frame drops, stuttering, or reconnection are detected, strategies such as interpolation compensation, frame drop marking, and reconnection reset are used to ensure the continuity of the timeline; when the time at the camera end is unreliable, unified time synchronization at the access end (such as NTP synchronization) can be used to rewrite or correct the frame timestamps to ensure the consistency of the boundaries of subsequent time windows and the comparability of historical statistics.

[0052] Preset time window length (For example, 2s, 5s, or 10s, the specific timeframe can be determined based on the personnel density and computing power configuration of the scenario), and the window sliding step size can be set; simultaneously, the data retention strategy for subsequent processes can be set, namely the minimum cache duration and database write frequency required for the statistics of the first and second historical time windows. Window sliding step size It can be Non-overlapping windows, or A sliding window is used to improve sensitivity.

[0053] Based on timestamps, the continuous frame sequence is divided into several time windows. }, where the k-th time window can be represented as: .in, Indicates the starting time of the window division. A non-negative integer, representing the index of the time window. Frames falling within the same time range of the window are grouped into a set of frames for that window. It generates window-level metadata for each window, including at least: window ID, start and end time, camera ID, number of frames within the window, and percentage of valid frames (the percentage after removing abnormal frames).

[0054] A quality metric is calculated for each time window, which can be based on a threshold of the number of valid frames within the window. When the quality metric falls below a preset threshold, the window is marked as a low-confidence window, and the reason is recorded (too many dropped frames, blurry image, black screen, etc.). This allows for weight reduction or skipping in subsequent target tracking and risk scoring stages, thereby reducing false alarms caused by low-quality data.

[0055] The windowed data is output in a standard structure for subsequent use by the S200 human detection and multi-target tracking. Simultaneously, the frame index corresponding to the window (e.g., first frame / middle frame / last frame) is retained to support rapid backtracking and evidence retrieval when combined with corresponding image fragments and sent to the monitoring terminal for alerts. The windowed data includes a time window identifier, camera identifier, window start and end time interval (using a left-closed, right-open format), and frame set. Window metadata (number of frames in the window, percentage of valid frames, frame drop rate, anomaly markers, etc.).

[0056] S200. Perform human target detection and multi-target tracking on each frame of the video stream within the current time window to obtain the unique tracking ID of the human target.

[0057] Time window output from S100 Read the frame sequence within the window The system reads windowed data (camera ID, timestamp range, frame rate, etc.); initializes the tracker state, including trajectory cache, target state vector (position, velocity, etc.), and ID counter (used to generate new target IDs), and can load scene calibration information (e.g., activity area ROI, entrance / exit area, desktop occlusion area, etc.). The tracker is a multi-target tracking module used to associate data with the same human target within adjacent time windows and output a unique tracking identifier. It maintains the target trajectory state and updates it based on the matching results. Based on human detection results, the tracker employs a data association strategy that includes motion consistency constraints and / or appearance similarity constraints, maintains the target trajectory, and assigns a consistent and unique tracking ID to the same human target; it retains the ID when the target is briefly lost and restores the association after the target is detected again.

[0058] Preprocessing is performed on each frame within the window to improve detection stability. Preprocessing operations include: scaling the size to the model input size, denoising / deblurring enhancement, brightness normalization, distortion correction for wide-angle lenses, and cropping the Region of Interest (ROI) (detection is only performed in the effective area to reduce false detections and computational cost). Abnormal frames (black screen, severe blur) are also marked and can be skipped or have their weight reduced in subsequent processing.

[0059] Perform human detection on each frame of the image within the window and output a set of human detection bounding boxes: ;in, Let J be the bounding box coordinates of the j-th person in the t-th frame. This indicates the confidence level. Human detection can be performed using the YOLO series. First, the model is loaded and parameters are configured, specifically the target category and the confidence threshold. (Range 0.3~0.7), overlap suppression threshold (For NMS, range 0.3–0.6), input size, then perform model input preprocessing on the t-th frame image, and input it into the human detection model to obtain the original candidate output set. (Contains a large number of candidate boxes): ;in, This represents the coordinates of the j-th candidate box. For candidate confidence, For category labels.

[0060] from Filter candidate boxes with the category "human body" and remove those below the confidence threshold. Candidates are selected. Non-Maximum Suppression (NMS) is performed on the removed candidate output set to remove duplicate bounding boxes generated from the same human body. NMS sorts the boxes from highest to lowest confidence, retaining the highest-scoring boxes and removing those with an overlap (IoU) exceeding the overlap suppression threshold. The other bounding boxes are then processed to obtain a deduplicated set of bounding boxes. This set is then subjected to regular filtering to remove boxes that clearly do not conform to human geometry or scene constraints. Rules include area filtering, aspect ratio filtering, and landing point filtering, resulting in the final set of human detection bounding boxes. .

[0061] For human detection box set For each candidate bounding box, extract target representation information for cross-frame association, including at least: geometric features (box center point, width, height, area, aspect ratio), appearance features, and human keypoint features. For example... Figure 7 As shown, after the image is divided into grids, it is processed by detection boxes, confidence scores, and class probabilities to obtain an output image with human targets.

[0062] For each trajectory that existed in the previous frame, perform state prediction to obtain the predicted position of that trajectory in the current frame. (Using Kalman filtering or a uniform velocity model to predict the center and scale of the bounding boxes) to match the current frame's detection box set; at the same time, update the trajectory survival timer.

[0063] Then, a cost matrix is ​​constructed and matching is performed to associate the "current frame detection box" with the "existing trajectory". The cost can be summarized by the following items:

[0064] Motion consistency cost: based on the IoU or center distance between the predicted bounding box and the detected bounding box;

[0065] Appearance similarity cost: embedding cosine distance / Euclidean distance;

[0066] The cost of pose consistency: key point similarity or orientation consistency.

[0067] Taking motion consistency cost and appearance similarity cost as examples, a comprehensive cost is constructed. :

[0068] ;

[0069] in, , Indicates the weighting coefficient; Indicates intersection, union, and ratio; This represents the predicted bounding box of the target position obtained from the motion model of the i-th historical trajectory at time t. This represents the j-th detection bounding box output by the human detection model at time t. This represents the appearance feature vector associated with the i-th trajectory (e.g., the embedding feature extracted from the human body bounding box region, which can be the average feature of the historical features of this trajectory). This represents the embedding feature vector extracted from the image region corresponding to the j-th detection box in frame t; This represents the feature distance metric function.

[0070] The optimal matching can be found using either the Hungarian algorithm or a greedy strategy. Three types of results are obtained after matching:

[0071] 1) Matched trajectory (updateable); 2) Unmatched detection (new target candidate); 3) Unmatched trajectory (may be occluded / departed).

[0072] For matched trajectories: Update the trajectory status (position, velocity, scale) with the corresponding detection box, and refresh the last appearance timestamp of the trajectory. For unmatched detection boxes: If the creation conditions (confidence, number of consecutive occurrences, being in a valid region, etc.) are met, create a new trajectory and assign a new unique tracking ID, recording its first appearance time. For unmatched trajectories: Enter a "short-term loss" state, allowing continued prediction within a preset number of lost frames L; if no match is found after exceeding the threshold, the trajectory is terminated and archived. L ranges from 3 to 30 frames, preferably 5 to 15 frames, or is equivalent to a maximum loss duration of 0.2 to 2.0 seconds.

[0073] To avoid frequent ID switching caused by occlusion, the following stabilization strategy can be introduced:

[0074] ID confirmation mechanism: A new trajectory must hit M consecutive frames before it is "confirmed and output", avoiding false IDs generated by noise false detection; M≤L;

[0075] Occlusion retention mechanism: Even if the trajectory is lost for a short time, the ID is retained and prediction continues to avoid being reassigned a new ID after reappearance due to occlusion;

[0076] Re-identification mechanism: When the trajectory terminates or is lost for a long time, if a highly similar target appears within a certain time / area, the original ID can be recovered or an ID mapping table can be established to enhance cross-occlusion consistency.

[0077] Within the current time window, output the tracking results for each frame or the summation of the window. Each frame's tracking result includes a unique tracking ID, timestamp, bounding box, confidence score, and / or a feature vector composed of appearance feature vectors and human keypoint features. The summation of the window's tracking results includes each person's ID within that time window, their trajectory within the window (the sequence of bounding boxes for each frame), and their first and last appearance times.

[0078] S300. Retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, take it as the first priority, and sort the rest in ascending order according to the spatial distance from the first priority.

[0079] Read the current time window The multi-target tracking output obtains the set of unique tracking IDs appearing within the current time window and the position of each unique tracking ID in the ground coordinate system. The location can be the representative location of the ID within the window (e.g., the average / median / last frame value of ground coordinates across multiple frames within the window), or it can include the ground trajectory sequence of the ID within the window for more precise distance / contact determination. Each tracking ID is a human target with the unique ID id within the current time window. Within the ground coordinate system (site plane coordinate system), the horizontal and vertical coordinate values ​​represent the position. The ground coordinate system is a two-dimensional coordinate system established with the site ground as the reference. The X-axis and Y-axis correspond to two mutually perpendicular directions in the site plane (e.g., along the length and width of the site), and the coordinate unit can be meters or centimeters (or normalized units).

[0080] Based on the preset first historical time window length Determine the time range of the first historical time window:

[0081] ;in, This is the start time of the current time window.

[0082] The database query criteria should include at least: location ID / camera ID and timestamp range. The system includes fields for marking event type and unique tracking ID. It retrieves a set of unique tracking IDs marked within the first historical time window from a pre-set database, and optionally reads the marking attributes associated with each ID (e.g., number of markings, last marking time, event type, and confidence level) as a basis for subsequent priority refinement.

[0083] The set of unique tracking IDs marked within the first historical time window is matched with the set of unique tracking IDs appearing in the current time window to obtain the intersection P. If the intersection P is not empty, it is determined that there exists a unique tracking ID that was marked within the first historical time window and appears in the current time window; if it is empty, the default sorting strategy for targets without historical marking is adopted.

[0084] When the intersection P is not empty, if the intersection P contains only one unique tracking ID, that unique tracking ID is directly taken as the first priority target; if the intersection P contains multiple unique tracking IDs, any of the following priority rules can be used to determine the first priority set:

[0085] 1. Select in descending order of the number of times the marker is placed within the first historical time window;

[0086] 2. Prioritize those with the most recent marking time.

[0087] The remaining unique tracking IDs in the set of unique tracking IDs appearing in the current time window are used to calculate the spatial distance d(id) between them and the first priority target based on the ground coordinate system (using Euclidean distance).

[0088] Then, construct the processing queue Q: place the first priority set at the front of the queue, and append the remaining unique tracking IDs in ascending order of their spatial distance from the first priority target.

[0089] When the intersection P is empty, to ensure that the queue can still be generated and used for subsequent steps, any of the following default sorting strategies can be used:

[0090] 1. Sort by ascending distance from the ground center of each unique tracking ID;

[0091] 2. Sort by the distance from each unique tracking ID to key areas of the venue (entrances / exits, service counters, ventilation openings, etc.) in ascending order;

[0092] 3. Sort by duration of continuous occurrence in descending order.

[0093] Output priority queue Q and its corresponding ground coordinates. The spatial distance d(id) serves as the input for subsequent "ROI division by priority, action sequence construction, event judgment and labeling into the database", thereby realizing the coordinated linkage of "historical labeled data - current occurrence data - ground spatial relationship data".

[0094] S400. Divide the human target into regions of interest in order of priority and construct the action sequence of the region of interest. Extract features from the action sequence and determine whether it meets the characteristics of the labeled event. If it does, label the unique tracking ID of the human target and update it to the database.

[0095] Among them, such as Figure 2 As shown, the specific steps for extracting features from an action sequence and determining whether it meets the characteristics of a labeled event are as follows:

[0096] S401. Extract features from the action sequence to obtain the torso tilt angle, head movement amplitude, and relative distance from key hand points to the mouth and nose area, and determine whether the corresponding feature thresholds are met.

[0097] S402. If so, perform event start and end location on the action sequence to obtain the start and end times of the marked events, and calculate the event confidence level;

[0098] S403. When the confidence level of an event is not lower than the preset event threshold, it is determined that the event characteristics are met.

[0099] Get the current time window The tracking results of each unique tracking ID are processed sequentially based on the priority queue Q, and a set of frame indexes (the set of all frame numbers in which the unique tracking ID appears in the window) is built for each unique tracking ID.

[0100] For each frame of image with a unique tracking ID within the window, the Region of Interest (ROI) is divided based on the human detection bounding box:

[0101] Trunk ROI: Covers the shoulder-hip area and is used to calculate the trunk forward tilt angle;

[0102] Head ROI: Covers the head region and is used to calculate the range of head movements;

[0103] Hand ROI: Covers the area near key points of both hands;

[0104] Mouth and nose ROI: Covers the mouth and nose area (which can be estimated from facial key points, nose tip / corner of mouth key points, or the mouth and nose position within the head ROI).

[0105] The ROIs are then scaled and positionally smoothed to reduce the impact of jitter on sequence features. Based on consecutive frames within the current time window and referencing human keypoints, the aforementioned ROIs are concatenated in frame order along the time dimension to form an action sequence. The relevant parameters in the action sequence are calculated as follows:

[0106] Calculate the torso tilt angle for each frame First, obtain the torso vector. , The angle between the angle and the vertical direction or the ground normal projection direction is calculated. ;in, Center point of the shoulder The hip center point is used as the output; the maximum anteversion angle within the current time window is output. .

[0107] Calculate the range of head movement based on key head points (e.g., top of head / tip of nose / center of head). Specifically, it is the maximum distance between any two header key points within the current time window (obtained by calculating and normalizing Euclidean distance).

[0108] Based on key points of the left and right hands and reference points of the mouth and nose Calculate the relative distance between the center points of the mouth and nose in each frame:

[0109] ;in, Represents the distance norm; This indicates the location of the key points on the left hand of the human body in frame t; The key points of the human right hand are located. The distance can be normalized (e.g., divided by head width / shoulder width) to obtain a normalized distance, and then the minimum distance within the current time window can be output. .

[0110] The above features are compared with a preset threshold to determine whether a candidate labeling event is triggered. Specifically, if... , , If the event starts and ends, proceed to the event start / end location step; otherwise, determine that the unique tracking ID in the current window does not meet the event marking characteristics and process the next unique tracking ID.

[0111] in, The forward tilt angle threshold is set in the range of 15° to 50°. Since coughing / sneezing is often accompanied by short-term obvious forward tilting, 20° to 35° is preferred. The threshold for head movement amplitude is set in the range of 0.05~0.40, preferably 0.10~0.25. The threshold value for the distance from the hand to the mouth and nose is set within the range of 0.10 to 0.60, preferably 0.15 to 0.35, since the distance between the hand and the mouth and nose is usually less than a certain ratio of the head width to the shoulder width.

[0112] Frames that meet the threshold condition are recorded as trigger frames, resulting in a trigger sequence. The trigger frames are then aggregated according to temporal continuity to form one or more candidate event segments. And impose a minimum duration constraint on the event segment (e.g., the event segment lasts for ≥ 1000 frames). (Only then is it considered effective) to remove brief noise segments. These are the start and end frame times of the m-th candidate event segment, respectively. The setting range is 5 to 10 frames.

[0113] exist Backtracking several frames, the frame in which the feature first consistently satisfies the condition is taken as the start time; Extending backwards for several frames, the frame where the feature continuously satisfies the condition is taken as the end time. Thus, the start and end times of the final event are obtained.

[0114] The features of each frame within the located event segment are fused to calculate the event confidence score. :

[0115] ;

[0116] in, , , These are the weighting coefficients; It is a normalization function, normalized to [0,1].

[0117] when If the action sequence satisfies the characteristics of a labeled event, then it is determined that the action sequence meets the criteria for a labeled event; otherwise, it is not labeled. The preset event threshold is set to achieve a relatively stable balance between reducing false alarms and maintaining sensitivity. The value range is 0.65 to 0.85, and 0.75 is preferred to meet the application scenario of this embodiment.

[0118] This application first extracts the torso tilt angle, head movement amplitude, and relative distance from the hand to the mouth and nose from the action sequence and performs a joint threshold judgment. This ensures that the system only triggers candidate events when the multi-dimensional behavioral evidence is consistent, effectively suppressing false triggers caused by single features such as occlusion, posture shaking, and nodding while speaking. After triggering, the start and end times of the event are located to obtain the start / end times and form an evidence time period, which is convenient for subsequent database statistics and risk accumulation, as well as for reviewing video clips for verification. Finally, the final judgment of event confidence and event threshold further filters out short-lived noise segments and weak evidence segments, realizing a closed-loop labeling system where only strong evidence is included in the database and can be located and played back. This improves the quality of labeled data and lays a more accurate foundation for subsequent frequency statistics, risk scoring, and graded early warning based on historical windows.

[0119] When determining whether the torso tilt angle, head movement amplitude, and relative distance from hand key points to the mouth and nose area meet the corresponding feature thresholds, if there are missing torso tilt angle, head movement amplitude, or relative distance from hand key points to the mouth and nose area, the number of effective key points and the visibility rate of key parts are obtained by comparing the pre-stored standard posture sequence and action sequence. The observability is then calculated by combining the target scale in the action sequence and added to the feature threshold judgment.

[0120] Specifically, if one of the following is not identified within the current time window: the torso tilt angle, the head movement amplitude, or the relative distance from the hand key point to the mouth and nose area, then the standard posture sequence (also known as the template sequence) corresponding to the marked event is retrieved from the pre-configured standard library. The standard posture sequence contains a set of standard key points of key parts and their temporal change patterns, including at least the visibility expectation and relative positional relationship range of key points such as the torso, head, hand, and mouth and nose area in typical events.

[0121] For each frame of the action sequence, the set of detected keypoints is extracted and compared with the set of standard keypoints in the standard pose sequence at the corresponding time phase (the phase is determined by alignment methods such as Dynamic Time Warping (DTW)). This yields the set of valid keypoints matched in each frame and the size of the set of keypoints that should appear in each frame. Within the current time window, the ratio of the number of valid keypoints to the number of keypoints that should appear is calculated to obtain the percentage of valid keypoints, which reflects the completeness of keypoint acquisition.

[0122] For the set of key body parts {torso, head, left hand, right hand, mouth and nose}, the visibility rate is calculated separately. Specifically, an indicator function (0 / 1 variable) is used to calculate the visibility rate of each body part: if the key point exists and the confidence level is greater than or equal to a threshold, it is set to 1; otherwise, it is set to 0. Finally, the visibility rates of all key body parts are averaged to obtain the overall visibility rate. The width, height, or area of ​​the human body detection bounding box is used as the target scale. The proportion of effective key points Visibility of key parts With target scale Fusion yields observability (0~1), the calculation formula is:

[0123] ;

[0124] in, , , These are the weighting coefficients; This is the scaling normalization function; This indicates that the interval is truncated to [0,1].

[0125] Set the minimum observability threshold Its value ranges from 0.4 to 0.6. When a candidate event is not triggered, or the decision is delayed until more frames have passed; when When, for feature threshold ( , , The correction is made using the following formula: ; The feature threshold before correction. The corrected feature threshold, To adjust the coefficient for the intensity of observability influence, a non-negative real number is used.

[0126] When the event threshold for the action sequence is finally satisfied... When a unique tracking ID is triggered, a tag record is generated corresponding to that unique tracking ID. The tag record includes at least the unique tracking ID, the start and end times of the tagged event, the event confidence level, the camera or location identifier, and the representative ground coordinates of the event occurrence. To make the tags interpretable and easy for the monitoring end to verify, an evidence fragment index is generated for the tag record and bound to it. The tag record is written to the database, and an index field is created for quick retrieval by time window or unique tracking ID. If the same unique tracking ID triggers the same tagged event consecutively within a short period, a merging strategy can be implemented: a merging time interval threshold is set, for example, 0.5s to 2s. If the start time of the new event and the end time interval of the most recent similar event in the database for that unique tracking ID are less than the merging time interval threshold, the two segments are merged into one event, and the end time is updated to the later one, the confidence level is updated with the maximum value or weighted by duration, the evidence fragments are merged, or the keyframe list is updated; otherwise, a new event detail record is added.

[0127] To support priority retrieval of the first historical time window and statistics of the second historical time window, it is preferable to maintain a tag index table (or status table) and update it after each tagging. When the tracker determines that the trajectory of a unique tracking ID has terminated (departure / long occlusion), the tagging status of that unique tracking ID can be set to historical status, but its event details and statistical information are still retained for historical retrieval when it reappears. After the database update is completed, it returns to the write status; if a write failure or network anomaly occurs, the tagging record is temporarily stored in a local queue and the write is retried, while the record is marked as "pending synchronization" to avoid losing critical tagging data.

[0128] This application addresses the common issue of missing or occluded keypoints in scenarios involving chess tables or crowds, by introducing observability performance to avoid misjudging events when information is insufficient, thereby improving the quality of labeled data. By fusing the effective keypoint count / visibility obtained through standard pose sequence comparison with the target scale, it can simultaneously reflect two types of unobservable causes: structural defects and insufficient resolution. By incorporating observability into threshold judgment, it achieves data quality-driven gating, reducing false alarms and improving the stability of event judgment.

[0129] S500. Based on the unique tracking ID marked in the current time window, count the number of times it was marked in the second historical time window from the updated database and sort it in descending order. Filter out all unique tracking IDs that exceed the safety threshold, analyze their behavioral characteristics one by one in descending order, calculate the risk score, match the corresponding risk level, and send the warning to the monitoring terminal in combination with the corresponding image fragment; wherein, the second historical time window is larger than the first historical time window.

[0130] Based on the current time window The set of unique tracking IDs that meet the characteristics of the marked event and have been written to the database, based on the preset second historical time window length. ( ), determine the time range of the second historical time window: ;in, This represents the end time of the current time window.

[0131] For each unique tracking ID located within the set, retrieve its information from the event details table or tag index table in the updated database within the second historical time window. The number of times an item is marked is sorted in descending order to obtain a sorted list. This sorting is used to prioritize the analysis of objects with "high-frequency triggering of marking events" to improve early warning efficiency and reduce unnecessary calculations. All unique tracking IDs exceeding the safety threshold are filtered out, and each unique tracking ID is processed one by one in descending order of the sorted list. Its value is retrieved from the database in the second historical window. Behavioral evidence within the data is used to form an analysis data package for risk assessment, including event type, start and end times, confidence level (Conf), duration, and evidence fragment index. The safety threshold, also known as the number of times an event is marked, ranges from 3 to 6 times; in this embodiment, 4 times is preferred. The calculation of a single risk score involves weighting relevant behavioral characteristics in the risk data package, such as the number of marked events and duration, and then correcting for observability. Observability is used to reduce weighting and avoid false high-risk errors caused by occlusion. The risk scores of unique tracking IDs exceeding the safety threshold within the current time window are averaged to obtain the total risk score R. In practical applications, the risk score with the highest single risk score within the current time window can also be selected as the risk score for that time window. If the safety threshold is not exceeded, no warning is triggered. When judging based on a single risk score, depending on the actual application needs, a warning can be issued for the first unique tracking ID exceeding the safety threshold, and subsequent warnings can be updated only when higher-level IDs are found.

[0132] Mapping the risk score R to the corresponding risk level, taking low / medium / high as an example, the mapping formula is as follows:

[0133] ;

[0134] in, Indicates the risk level of the unique tracking ID; , , This indicates the first risk level, the second risk level, and the third risk level; , This indicates the risk classification threshold.

[0135] Construct an alert data packet for the unique tracking ID that meets the alert criteria in the current time window, including at least:

[0136] The system uniquely tracks the ID, risk score, risk level, and number of times the event is marked in the second historical window; it also tracks the keyframe of the highest confidence event or the video clip of the event that contributes the most to the risk score within the second historical window. The alert package is sent to the monitoring terminal, and the sending time, receipt confirmation, and handling status are recorded. Once the monitoring terminal confirms or processes the alert, the handling result can be written back to the database as closed-loop data for subsequent threshold adaptation and model calibration.

[0137] Therefore, this embodiment achieves continuous identification and tracing at the individual level. It utilizes the marked IDs within the first historical time window as the first priority, and sorts other targets according to their spatial distance. It prioritizes ROI and action sequence analysis for high-risk associated objects, reducing the computational cost of analyzing all individuals equally and improving real-time early warning capabilities. Based on the second historical time window (larger than the first), it counts the number of times objects are marked and compares this count with a safety threshold to filter out objects that frequently trigger risk events. Then, it calculates risk scores based on behavioral characteristics and matches risk levels, achieving an upgrade from event detection to risk assessment and tiered handling. By prioritizing historically high-risk individuals and sorting by spatial distance, the system focuses more on individuals with potentially higher transmission ties and their neighboring populations, facilitating earlier detection of potential cluster transmission risks and timely early warning.

[0138] The main scheme of this application has been described above, such as... Figure 3 As shown, the following further explains the specific steps for calculating risk scores by combining them with the surrounding population, that is, analyzing behavioral characteristics one by one in descending order and comprehensively calculating the risk score:

[0139] S501. After sorting each unique tracking ID, a unified timestamp is applied to the marked event records within the second historical time window, and the cumulative duration of the marked events for each unique tracking ID is calculated.

[0140] S502. Using the unique tracking ID as the center, divide the contact area, and treat other human targets detected within the contact area as the surrounding population, calculate the density value and contact duration of the surrounding population in the contact area;

[0141] S503. Using the unique tracking ID as an index, construct a behavioral feature vector by the number of marked events, cumulative duration, density value of the contact area, and contact duration, and calculate the single risk score for each unique tracking ID after normalization and weighting.

[0142] S504. Perform cluster analysis on the set of behavioral feature vectors to obtain at least one behavioral cluster, and correct the single risk score of each unique tracking ID based on the intra-cluster statistics of the behavioral clusters;

[0143] S505. Merge all the single risk scores corrected by the unique tracking IDs to obtain the risk score.

[0144] Specifically, for each unique tracking ID after sorting, retrieve its value from the database within the second historical time window. The set of tagged events, each event record must contain at least: event start time, end time, event type, and event confidence level.

[0145] Unify the start and end times of events to the same time base, such as the system's unified UTC time or local time, and align them with uniform granularity. Calculate the cumulative duration based on the start and end times of the marked events. Obtain the unique tracking ID and its position within the second historical time window from the multi-target tracking results. , The contact area is constructed centered on the ground location of the unique tracking ID. Preferably, a radius of [missing information] is used. The circular region. The calculation formula is:

[0146] Among them, among them It can be 1m to 2m in diameter to meet the risk of close contact.

[0147] in, These are the horizontal and vertical coordinates of the unique tracking ID on the ground at time t, respectively. is the position vector of other human targets in the ground coordinate system.

[0148] For each time t, obtain the ground coordinates of all other human targets under the same timestamp alignment. If satisfied The target j is included in the set of people surrounding the unique tracking ID at that moment. Then, the density value of the contact area is calculated using the following formula:

[0149] ;

[0150] in, This represents the surrounding density value of the unique tracking ID at time t; This refers to all other human targets besides the unique tracking ID being processed; This represents the ground coordinates of the unique tracking ID at time t; This represents the action term of the kernel function on the distance; Indicates the indicator function (0 / 1 function).

[0151] Regarding the calculation of contact duration, a contact indicator variable is defined as follows: The formula for calculating the contact duration is: ;in, Indicates a unique tracking ID in the second historical time window Duration of contact within; Indicates the time step.

[0152] Use the current unique tracking ID in The number of marked events, cumulative duration, density value of the contact area, and contact duration constitute the behavioral feature vector. Each dimension of the feature vector is normalized, and then weighted according to the vector... Calculate the individual risk score. Perform clustering on all unique tracking IDs to be evaluated to obtain at least one behavioral cluster. The clustering method can be K-means, GMM, or hierarchical clustering. For each behavioral cluster, calculate the within-cluster statistics, which can be the mean and standard deviation of the individual risk score within the cluster, or the within-cluster quantiles.

[0153] Each unique tracking ID is adjusted for a single risk score by combining the statistics of its cluster. The adjustment formula is as follows: ;

[0154] in, This is a correction factor, with a value range of 0.2–1.0; This indicates the corrected single-risk score; This indicates the single-risk score before correction; This represents the unique tracking ID combined with the mean single-risk score of its cluster.

[0155] The risk score R is obtained by averaging all individual risk scores. Alternatively, the maximum individual risk score within the current time window can be used as the risk score R for that time window.

[0156] This scheme uses a second historical time window as a statistical scale, and jointly models event behavior evidence (number of events and cumulative duration) and spatial contact evidence (density of contact area and duration of contact) on the same time benchmark to form a quantifiable behavioral feature vector and calculate a single risk score. This transforms the randomness of a single trigger into a cumulative and comparable risk quantity, which can reflect the high-frequency abnormal behavior of the target individual over a long period of time and correlate it with the exposure intensity of the surrounding population. This improves the interpretability and stability of infection risk assessment and provides a more reliable basis for subsequent graded early warning.

[0157] like Figure 4 As shown below, the specific steps for merging the single risk scores after correcting all unique tracking IDs are further given:

[0158] S511. Calculate the Top-K mean from the corrected single-risk scores;

[0159] S512. From the corrected single-risk scores, select scores that are greater than the first risk threshold and scores that are in the range between the first risk threshold and the second risk threshold, calculate their averages, and then weight them to obtain the risk mean.

[0160] S513. The risk score is obtained by merging the total mean of all corrected single risk scores with the Top-K mean and the risk mean.

[0161] Specifically, the corrected single-risk scores of all entities to be merged are aggregated and sorted in descending order to obtain the sequence: ; This represents the nth adjusted single-risk score. Let K be a preset integer, and 1 ≤ K ≤ n. Calculate the Top-K mean using the first K scores. This highlights the small group of individuals at the highest risk, who are sensitive to the risks of cluster outbreaks. Among them, Let i represent the i-th modified single-risk score from 1 to K.

[0162] Calculate the mean of all corrected single-risk scores from 1 to n. .

[0163] First risk threshold Second risk threshold Taking the fraction [0,1] as an example, The value range is [0.75, 0.9). The value range is [0.6, 0.75].

[0164] Selecting those with a risk score greater than the first risk threshold from the revised single-risk score. and located at the first risk threshold Second risk threshold The first risk mean is calculated by averaging the scores of each interval. Second risk mean The weighted average risk is obtained as follows: ;in, , The weighting coefficients are non-negative and for all... By assigning higher weights, the risk mean becomes more sensitive to high-risk groups.

[0165] Final risk score R: ;

[0166] , , Let be the weighting coefficients, and let sum to 1. The range is [0.15, 0.30]. The range is [0.30, 0.50]. The range is [0.25, 0.45].

[0167] By simultaneously incorporating three statistical measures—the overall mean, the Top-K mean, and the risk mean—the corrected single-risk score is aggregated across multiple scales. This allows the final risk score to reflect the overall risk profile of the population (smoothing effect provided by the overall mean), maintain sufficient sensitivity to a small number of extremely high-risk individuals (the Top-K mean highlights the high-risk tails), and characterize the structural changes in individuals within the medium-to-high-risk range (risk means segmented based on the first and second risk thresholds). Therefore, compared to simply taking the maximum value or the average value, this method strikes a balance between avoiding fluctuations caused by extreme values ​​and avoiding underreporting due to average dilution, thus improving the stability and sensitivity of risk assessment.

[0168] To facilitate understanding of the above embodiments, a specific application scenario of the above embodiments will be used as an example for illustration below:

[0169] Taking a community senior activity center as an example, suppose cameras are installed inside the center, monitoring areas including the chess and card area, dance area, and rest area. The cameras capture video streams in real time, and the system extracts one frame every 5 seconds for analysis. It identifies seniors and tracks their behavior, such as the angle of their torso leaning forward and the distance between their hands and their mouth and nose. Suppose six human targets are detected in the current time window:

[0170] Target A: Sit at a table, maintaining a distance of approximately 1 meter from targets B, C, and D for 20 minutes.

[0171] Target B: Frequent contact with targets A, C, D, and E for 15 minutes.

[0172] Objective F: Maintain a distance from others for 5 minutes.

[0173] For target A, the system detected that its torso leans forward at an angle of 35°, its head movement range is 0.20 meters, and its hand touches its mouth and nose at a distance of 0.30 meters (consistent with high-risk behavior).

[0174] For target B, the system detected that its torso leaned forward at an angle of 40°, its head movement range was 0.25 meters, and its hand touched its mouth and nose at a distance of 0.35 meters (consistent with high-risk behavior).

[0175] For target C, the system detected that the torso tilted forward at an angle of 10°, the head movement range was 0.05 meters, and the hands did not touch the mouth and nose, which is considered a low-risk behavior.

[0176] For target D, the system detected that its torso tilted forward at an angle of 10°, its head movement range was 0.05 meters, and its hands did not touch its mouth or nose, which is considered a low-risk behavior.

[0177] For target E, the system detected that its torso tilted forward at an angle of 5°, its head movement range was 0.08 meters, and its hands did not touch its mouth or nose, which is considered a low-risk behavior.

[0178] For target F, the system detected that its torso tilted forward at an angle of 10°, its head movement range was 0.08 meters, and its hands did not touch its mouth or nose, which is considered a low-risk behavior.

[0179] Target A's action sequence was identified as high-risk behavior within a 5-second time window, and its hands approached its mouth and nose more than twice, meeting the event characteristics. Target B also exhibited similar behavior, with its hands touching its mouth and nose more than three times.

[0180] Assuming a safety threshold of 3, only targets A and B exceed the threshold. After normalizing their behavioral feature vectors, their risk scores are calculated to be 10.43 and 8.44, respectively. The result after fusion calculation is greater than... (Taking 0.8 as an example in this embodiment), the current risk level is determined to be the first level, triggering an alarm. The corresponding image segment, time window information, and risk level are sent to the monitoring terminal. Relevant management personnel can manually review the data sent by the monitoring terminal and perform corresponding operations. The location relationship of the target can be referenced... Figure 8 As shown, the coordinate system is a ground coordinate system, with the X-axis and Y-axis corresponding to two mutually perpendicular directions in the plane of the site (e.g., along the length and width of the site), and the coordinate unit can be meters.

[0181] The main scheme of this application has been introduced above. The following details the calculation of the density and duration of contact with surrounding people in the contact area, and also includes the calculation of the mask risk index within the contact area. For example... Figure 5 As shown, the specific steps are as follows:

[0182] S521. Within the contact area, perform mask-wearing status recognition on the surrounding population and maintain a mask-wearing status machine to obtain the corresponding mask status sequence;

[0183] S522. Calculate the single mask exposure index of the surrounding population based on the mask status sequence, and merge all single mask exposure indices of the surrounding population to obtain the mask risk index in the contact area;

[0184] S523. Correct the single risk score of each unique tracking ID by combining the mask risk index with the intra-cluster statistics of the behavioral cluster.

[0185] Taking a human target with a unique tracking ID of id as an example, in the second historical time window Contact area constructed within the target area and the surrounding crowd By using a classifier, data is collected from the surrounding population. Perform mask-wearing status recognition in the face / head ROI and output discrete states. and confidence level In practical applications, the YOLO model can also be used as the classifier. The mask status includes at least: Not wearing; Wearing instructions (covering the mouth and nose); Improper wearing (exposing nose / mouth / hanging on chin); Undecidable / Occlusion. A mask-wearing state machine is maintained for each surrounding target j to stabilize the frame-by-frame classification results. Specifically, the discrete states of each frame are input. and confidence level Output stable state If confidence level If the value is less than a preset threshold (range 0.6-0.8), the frame is recorded as "undeterminable" and no state transition is triggered. Only when a certain state occurs consecutively for at least [number missing] [time missing] will the frame be considered "undeterminable" and no state transition will occur. The system switches to this state only when a specific frame is reached; different confirmation frame numbers can be set for switching from "proper wearing" to "not wearing / inappropriate wearing" and vice versa to avoid frequent jitter. This yields the mask status sequence of the surrounding target within the time window. The frame rate can be set from 2 to 15, with 3 to 6 frames being the preferred setting.

[0186] The classifier mainly consists of the following modules:

[0187] Input preprocessing layer: The input image is a cropped image of a face or head region, with a size of 224×224 or 256×256 pixels, and is input to the next layer after data augmentation.

[0188] Feature extraction layer: ResNet-50 is used to extract low-level and high-level features (such as edges, textures, shapes, etc.) from the input image. Max pooling is used to reduce the spatial dimension and computational cost, while retaining the most salient features, and the features are then input to the next layer.

[0189] Fully connected layer and classification layer: After extracting features, multiple fully connected layers are used to further fuse these features, and finally output the prediction result of mask wearing status. In the last layer, the Softmax activation function is used to normalize the probability of each category to the range of [0,1], which represents the probability value of each mask wearing status.

[0190] Output layer: The prediction results of the last layer are converted into four types of output, representing different mask wearing states, and the confidence level of each category (i.e. the probability value of each category).

[0191] Assign risk weights to different mask statuses, for example The weight can be 0. Assign the highest weight. Assign higher weights, The allocation is slightly lower than The weights are determined. For each surrounding target j, the single-mask exposure index is calculated based on its state sequence. :

[0192] ;

[0193] in, This represents the distance decay weighting function; This represents the set of valid time indices for the surrounding target j; This represents the stable mask state of the surrounding target j at time t; The function representing the risk weighting of mask status; This represents the ground distance between the surrounding target j and the central target id at time t, which is calculated from the Euclidean distance between their ground coordinates. It represents a very small positive number.

[0194] The mask risk index is calculated by aggregating the exposure indicators of individual masks within the contact area and using a density-weighted fusion method. :

[0195] ;

[0196] in, The cumulative weight of the contact time between surrounding targets and target ID.

[0197] Calculate the intra-cluster mean of the mask risk index for each behavioral cluster, and compare it with... Calculate deviation Based on deviation The single risk score is adjusted using the following formula: ; This is a correction factor; This indicates the corrected single-risk score; This represents the single-risk score before correction.

[0198] By introducing mask-wearing status recognition and state machine stabilization within the contact area, the mask risk index of the contact area is calculated and the single risk score is corrected by combining behavioral cluster statistics. This allows the risk assessment to consider both individual abnormal behavior and the surrounding protection level, improving the sensitivity to real exposure scenarios and reducing false alarms caused by momentary misidentification.

[0199] The above section detailed the steps for calculating the risk index of masks within the contact area. Below, after performing human target detection and multi-target tracking on each frame of the video stream within the current time window, a detailed explanation of risk identification based on environmental factors will be provided. Figure 6As shown, the specific steps are as follows:

[0200] S531. Obtain the date of the video stream captured within the current time window, and determine whether it falls within the preset date range. If so, determine that it is in the peak season for influenza.

[0201] S532. Based on all unique tracking IDs appearing within the current time window, perform mask wearing status identification and maintain the mask wearing state machine to obtain the corresponding mask status sequence;

[0202] S533. Calculate the proportion of people not wearing masks within the current time window based on the mask status sequence. If the proportion of people not wearing masks is higher than the preset mask threshold, it is determined that there is an exposure risk. The corresponding image segment is then sent to the monitoring terminal for early warning.

[0203] Specifically, the current time window is obtained from video stream metadata or system time. The shooting date is determined by the system time of the server / edge device. If the video stream has its own timestamp, the timestamp of the first or middle frame of the time window is used as the shooting date. If there is no metadata, the system time of the server / edge device is used as the shooting date. The set of dates for peak influenza seasons is read from the preset configuration table. If the shooting date falls within this range, it is determined that the shooting date is during a peak influenza season.

[0204] The system obtains a unique set of tracking IDs appearing in the current time window from the multi-target tracking output. It then retrieves the mask status and mask status sequence. The proportion of frames judged as not wearing masks out of the total number of valid frames in the mask status sequence is used as the "not wearing a mask ratio," which is compared with a mask threshold. If the not wearing mask ratio is higher than a preset mask threshold, an exposure risk is determined. The mask threshold ranges from 0.2 to 0.5.

[0205] Taking a mask threshold of 0.35 as an example, if 20 IDs appear within the current time window, only 18 IDs meet the valid frame count requirement (included in the statistics). Of these, 6 IDs are determined not to be wearing masks, resulting in a mask-less ratio of 6 / 18 = 0.333. Since this ratio does not exceed 0.35, it is determined that there is currently no exposure risk. If the mask threshold were 0.3, then the current ratio exceeds the threshold, and an exposure risk is determined. Evidence fragments are selected from the window, such as the frame with the most people not wearing masks, and the frame number / timestamp, camera ID, and location ID are recorded. Then, an early warning data packet is constructed, including the time window identifier, shooting date, mask-less ratio, and evidence fragment index. The early warning packet is sent to the monitoring terminal, and the early warning sending time, reception status, and handling status are recorded locally or in the database.

[0206] By performing mask recognition on all individuals within a time window and stabilizing it through a state machine, then using the proportion of individuals not wearing masks for gating judgment, and adjusting the threshold or risk level in conjunction with the date factor of the peak influenza season, rapid and interpretable early warning of environmental exposure risks can be achieved, reducing false alarms caused by single-frame misjudgment and improving the practicality of the early warning.

[0207] Example 2

[0208] like Figure 9 As shown in the figure, this embodiment introduces a multi-source behavioral data intelligent early warning system for community population infection risk, including a data acquisition module, a target tracking module, a priority setting module, an event judgment module, and a data output module.

[0209] The data acquisition module is used to acquire real-time video streams captured by cameras in the target public activity area and to divide the video streams into time windows;

[0210] The target tracking module is used to perform human target detection and multi-target tracking on each frame of the video stream within the current time window, and obtain a unique tracking ID for the human target;

[0211] The priority setting module is used to retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, it is set as the first priority, and the rest are sorted in ascending order according to the spatial distance from the first priority.

[0212] The event judgment module is used to divide the human target into regions of interest in order of priority and construct the action sequence of the regions of interest. The action sequence is then used to extract features and determine whether it meets the characteristics of the event to be marked. If it does, the unique tracking ID of the human target is marked and updated to the database.

[0213] The data output module is used to count the number of times a unique tracking ID is marked in the second historical time window from the updated database based on the unique tracking ID marked in the current time window, sort it in descending order, filter out all unique tracking IDs that exceed the safety threshold, analyze their behavioral characteristics one by one in descending order, calculate the risk score, match the corresponding risk level, and send the warning to the monitoring terminal in combination with the corresponding image fragment.

[0214] The second historical time window is larger than the first historical time window.

[0215] This embodiment has the same beneficial effects as Embodiment 1.

[0216] Example 3

[0217] This embodiment introduces a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned intelligent early warning method based on multi-source behavioral data for community population infection risk.

[0218] When applying the multi-source behavioral data-driven intelligent early warning method for community infection risk, it can be implemented in software form, such as a program designed to run independently on a computer-readable storage medium, which could be a USB flash drive or a USB security token. The program could be designed to start the entire method via an external trigger.

[0219] Example 4

[0220] This embodiment introduces an intelligent early warning device based on multi-source behavioral data for community infection risk, including at least one camera, memory, and processor.

[0221] The cameras are installed in the target public activity area to collect real-time video streams of the target public activity area;

[0222] The memory is used to store computer programs and a pre-set database, which at least stores unique tracking IDs and their tagging information within the first and second historical time windows.

[0223] The processor communicates with the camera and the memory. The processor is used to call and execute the computer program stored in the memory. When the program is executed by the processor, it implements the steps of the aforementioned intelligent early warning method for multi-source behavioral data on the risk of infection in the community population.

[0224] Cameras are preferably installed on the top or walls of entrances, corridors, chess and card areas, activity rooms, and other areas in community public activity spaces to form a field of view covering the areas where people are active. Single cameras or multi-camera networking can be used, and the fields of view of multiple cameras can partially overlap to reduce obstruction and blind spots.

[0225] The video stream captured by the camera is preferably at 1080p resolution, with a frame rate of 15–30fps; the video encoding can be H.264 / H.265. The video stream output by the camera carries timestamp information or is generated by the processor for unified time synchronization, which is used for subsequent time window division and historical time window statistics.

[0226] To obtain the position of a human target in the ground coordinate system, the camera can be pre-calibrated for homography or camera calibration to establish a mapping relationship between image coordinates and ground coordinates. The processor can then convert the target position into ground coordinates (e.g., in meters) based on this mapping for purposes such as contact area division and density calculation.

[0227] The memory stores computer programs used for human detection, multi-target tracking, posture / key point extraction, mask status recognition, event determination, risk scoring and early warning push; it can also store model parameter files, threshold configuration files, camera calibration parameters, and location area division parameters.

[0228] The processor can be deployed on edge computing devices (or central servers), or an edge + central collaborative architecture can be adopted: the edge side completes detection and tracking and initial event judgment, while the central side completes historical statistics, cluster correction, risk scoring, and alarm aggregation. Tasks such as video decoding, human detection / tracking, pose key point inference, mask recognition, feature calculation, database read / write, and early warning push can be divided into multiple threads / processes or pipeline modules to meet real-time requirements; when there are many cameras, they can be processed in parallel according to camera channels.

[0229] Tasks such as video decoding, human detection / tracking, posture key point inference, mask recognition, feature calculation, database reading and writing, and early warning push are divided into multiple threads / processes or pipeline modules to meet real-time requirements; when there are many cameras, they can be processed in parallel according to camera channels.

[0230] like Figure 10 As shown, if multiple cameras exist in an activity room, to avoid wasting computing power in practical applications, the camera with the widest field of view is designated as the main camera, and the others are used as auxiliary cameras. Clock synchronization of all cameras: Within the local area network, cameras can use NTP (Network Time Protocol) to ensure time synchronization of video streams, guaranteeing that the acquisition time of the main camera and auxiliary cameras is consistent. Auxiliary cameras synchronize their acquired video streams with the main camera to ensure the accuracy of event detection. Auxiliary cameras align their data with the main camera's video stream using timestamps, ensuring that the activity events from all cameras can be merged into unified tracking information.

[0231] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.

Claims

1. A multi-source behavioral data-driven intelligent early warning method for community population infection risk, characterized in that, Includes the following steps: Acquire real-time video streams captured by cameras within the target public activity area, and divide the video streams into time windows; Perform human target detection and multi-target tracking on each frame of the video stream within the current time window to obtain a unique tracking ID for the human target; Retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, take it as the first priority, and sort the rest in ascending order of spatial distance from the first priority; The human target is divided into regions of interest in order of priority and the action sequence of the region of interest is constructed. The action sequence is then subjected to feature extraction to determine whether it meets the characteristics of the labeled event. If it does, the unique tracking ID of the human target is labeled and updated to the database. Based on the unique tracking ID marked in the current time window, the number of times it was marked in the second historical time window is counted from the updated database and sorted in descending order. All unique tracking IDs that exceed the safety threshold are filtered out, and their behavioral characteristics are analyzed one by one in descending order. The risk score is calculated and matched with the corresponding risk level. The corresponding image fragment is then sent to the monitoring terminal for early warning. The second historical time window is larger than the first historical time window.

2. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 1, characterized in that, The specific steps for feature extraction from action sequences and subsequent determination of whether they meet the characteristics of a labeled event are as follows: Feature extraction is performed on the action sequence to obtain the torso tilt angle, head movement amplitude, and relative distance from key hand points to the mouth and nose area, and to determine whether the corresponding feature thresholds are met. If so, the start and end times of the action sequence are located to obtain the start and end times of the marked events, and the event confidence is calculated; When the confidence level of an event is not lower than the preset event threshold, it is determined that the event meets the characteristics of the event.

3. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 2, characterized in that, When determining whether the torso tilt angle, head movement amplitude, and relative distance from hand key points to the mouth and nose area meet the corresponding feature thresholds, if there are missing torso tilt angle, head movement amplitude, or relative distance from hand key points to the mouth and nose area, the number of effective key points and the visibility rate of key parts are obtained by comparing the pre-stored standard posture sequence and action sequence. The observability is then calculated by combining the target scale in the action sequence and added to the feature threshold judgment.

4. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 3, characterized in that, The specific steps for analyzing behavioral characteristics one by one in descending order and calculating the risk score are as follows: After sorting each unique tracking ID, a unified timestamp is applied to the marked event records within the second historical time window, and the cumulative duration of the marked events for each unique tracking ID is calculated. Using a unique tracking ID as the center, the contact area is divided. Other human targets detected within the contact area are considered as the surrounding population. The density value and contact duration of the surrounding population in the contact area are calculated. Using the unique tracking ID as an index, the number of marked events, cumulative duration, density value of the contact area, and contact duration are used to construct a behavioral feature vector. The single risk score of the unique tracking ID is calculated by normalization and weighting, and the risk score is obtained after correction.

5. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 4, characterized in that, The specific steps to obtain the risk score after correction are as follows: Cluster analysis is performed on the set of behavioral feature vectors to obtain at least one behavioral cluster, and the single risk score of each unique tracking ID is corrected based on the intra-cluster statistics of the behavioral clusters. The risk score is obtained by merging the individual risk scores corrected for all unique tracking IDs.

6. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 5, characterized in that, The specific steps for merging the single-risk scores corrected for all unique tracking IDs are as follows: The Top-K mean was calculated from the corrected single-risk score; The scores that are greater than the first risk threshold and those that are in the range between the first and second risk thresholds are selected from the corrected single risk scores. The average of these scores is calculated and then weighted to obtain the risk mean. The risk score is obtained by combining the overall mean of all corrected single-risk scores with the Top-K mean and the risk mean.

7. The intelligent early warning method for multi-source behavioral data on infection risk in community populations according to claim 1, characterized in that, After performing human target detection and multi-target tracking on each frame of the video stream within the current time window, the process also includes risk identification based on environmental factors. The specific steps are as follows: Obtain the date of the video stream captured within the current time window and determine whether it falls within a preset date range; if so, determine that it is during the peak flu season. Based on all unique tracking IDs appearing within the current time window, perform mask wearing status identification and maintain the mask wearing state machine to obtain the corresponding mask status sequence; The proportion of people not wearing masks within the current time window is calculated based on the mask status sequence. If the proportion of people not wearing masks is higher than the preset mask threshold, it is determined that there is an exposure risk, and an early warning is sent to the monitoring terminal in conjunction with the corresponding image segment.

8. A multi-source behavioral data-driven intelligent early warning system for community population infection risk, characterized in that, include: The data acquisition module is used to acquire real-time video streams captured by cameras in the target public activity area and to divide the video streams into time windows. The target tracking module is used to perform human target detection and multi-target tracking on each frame of the video stream within the current time window, and obtain a unique tracking ID for the human target. The priority setting module is used to retrieve the unique tracking ID marked in the first historical time window from the pre-set database, determine whether it appears in the video stream in the current time window, and if so, set it as the first priority, and sort the rest in ascending order according to the spatial distance from the first priority. The event judgment module is used to divide the human target into regions of interest in order of priority and construct the action sequence of the regions of interest. The action sequence is then used to extract features and determine whether it meets the characteristics of the event to be marked. If it does, the unique tracking ID of the human target is marked and updated to the database. The data output module is used to count the number of times a unique tracking ID is marked in the second historical time window from the updated database based on the unique tracking ID marked in the current time window, sort it in descending order, filter out all unique tracking IDs that exceed the safety threshold, analyze their behavioral characteristics one by one in descending order, calculate the risk score, match the corresponding risk level, and send the warning to the monitoring terminal in combination with the corresponding image fragment. The second historical time window is larger than the first historical time window.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the intelligent early warning method for multi-source behavioral data on infection risk in community populations as described in any one of claims 1 to 7.

10. A multi-source behavioral data-driven intelligent early warning device for community population infection risk, characterized in that, include: At least one camera, which is installed in the target public activity area, is used to collect real-time video streams of the target public activity area; A memory for storing computer programs and a pre-set database, the database being used to store at least a unique tracking ID and its tagging information within a first historical time window and a second historical time window; A processor, communicatively connected to the camera and the memory, is used to call and execute a computer program stored in the memory, which, when executed by the processor, implements the steps of the intelligent early warning method for multi-source behavioral data on community population infection risk as described in any one of claims 1 to 7.