Visual display method and system applied to security and protection and storage medium

By constructing a road coordinate system and fusing multimodal images, the problems of high information density and insufficient structured expression in existing security systems have been solved. This enables intelligent processing of surveillance images and real-time risk warnings, improving the display accuracy and response efficiency of security systems.

CN121009211AActive Publication Date: 2025-11-25JINGGANGSHAN YUJIE FIRE SCI & TECH

Patent Information

Application Number
CN202511097753.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-25
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing security systems suffer from high information density and a lack of structured abstraction methods at the display and interaction levels. This makes it difficult to quickly obtain the spatial location and event attributes of key alarms or potential hazard areas, and also lacks the ability to understand the structural semantics of target areas, making it difficult to meet the comprehensive judgment needs for passage safety in complex scenarios.

Method used

By acquiring data streams collected by security equipment, performing data compression and multimodal image fusion, identifying road boundary point sets and constructing a road coordinate system, detecting dynamic and static targets, analyzing target motion states, predicting the overlap time points of target motion states, and executing voice prompt tasks through IoT transmission.

Benefits of technology

It enables structured representation of surveillance footage and intelligent target recognition, improving the display accuracy, judgment depth, and response efficiency of security systems. It can proactively identify potential access risks and provide real-time alerts, enhancing security management capabilities in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009211A_ABST
    Figure CN121009211A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of security and protection visualization, in particular to a security and protection visualization display method and system and a storage medium. The method comprises the following steps: acquiring a data stream acquired by security and protection equipment, and compressing the data stream into a unified acquisition frame set; identifying a road boundary point set in the unified collection frame set, and detecting a dynamic target and a static target at each moment in the unified collection frame set; performing semantic annotation according to the motion state change condition of the dynamic target at each moment in the unified collection frame set to obtain a road boundary semantic distribution map; according to the road boundary semantic distribution map, identifying and predicting the motion state change condition of the static target; and predicting an overlapping time point of the motion state of the target based on the motion state change condition of the static target and the motion state change condition of the dynamic target, and transmitting the overlapping time point to communication equipment of the static target through the Internet of Things so as to execute a voice prompt task. According to the invention, the display precision and response efficiency of security and protection visualization can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the security visualization technical field, and particularly relates to a security visualization display method and system and a storage medium. BACKGROUND

[0002] In the prior art, security devices mainly rely on monitoring cameras to collect images of key areas, and realize image storage, playback and manual checking through a background system. Some systems introduce image recognition, abnormal behavior analysis and other auxiliary functions to improve the response capability to key events.

[0003] However, the existing security system still has great limitations in display and interaction. The video picture information density is large, and there is a lack of structured abstraction means, so that the value-keeping personnel cannot obtain the spatial position and event attribute of the key police situation or hidden danger area in a short time.

[0004] In addition, the current security visualization means mostly rely on two-dimensional images or simple interface superposition, and lack the ability to understand the structure semantics of the target area, and it is difficult to meet the comprehensive judgment demand of the safety of the complex scene. SUMMARY

[0005] Therefore, it is necessary to provide a security visualization display method and system and a storage medium to solve at least one of the above technical problems.

[0006] To achieve the above purpose, a security visualization display method comprises the following steps:

[0007] Step S1: acquiring security device collected data stream, and compressing the security device collected data stream to obtain a uniform collection frame set;

[0008] Step S2: identifying a road boundary point set in the uniform collection frame set, and taking the road boundary point set as the main axis of the road coordinate system, detecting dynamic targets and static targets at each time in the uniform collection frame set;

[0009] Step S3: according to the motion state change of the dynamic targets at each time in the uniform collection frame set, distinguishing the categories of the dynamic targets, and semantically labeling the categories of the dynamic targets to obtain a road boundary semantic distribution map;

[0010] Step S4: binding the corresponding image slices and time stamps of each node in the road boundary semantic distribution map to obtain a road boundary dynamic distribution map; according to the road boundary dynamic distribution map, analyzing the interaction of the dynamic targets and the static targets, and identifying and predicting the motion state change of the static targets;

[0011] Step S5: predicting an overlap time point of the target motion state based on the static target motion state change and the dynamic target motion state change, and transmitting the overlap time point to the communication device of the static target through the Internet of Things to perform a voice prompt task.

[0012] Optionally, step S1 comprises:

[0013] Step S11: obtaining security device collection data stream and device auxiliary state data through the security device deployed in the security area, wherein the security device collection data stream comprises RGB visible light image sequence, depth map sequence and infrared thermal imaging data;

[0014] Step S12: performing device timestamp alignment processing on the security device collection data stream, and performing frame reconstruction interpolation on the timestamp alignment processing result according to the camera time sequence in the device auxiliary state data to obtain a synchronous alignment image frame sequence;

[0015] Step S13: performing spatial view angle transformation processing on the synchronous alignment image frame sequence according to the pose parameters and lens intrinsic parameters of each collection device in the device auxiliary state data to generate a view angle normalized image set;

[0016] Step S14: performing inter-frame difference feature extraction on the view angle normalized image set, eliminating redundant frames without target motion change, and effectively sampling and compressing the remaining frames to generate a high-confidence frame sequence;

[0017] Step S15: data splicing and packaging the RGB visible light image, depth map, infrared thermal imaging and timestamp in the high-confidence frame sequence to obtain a unified collection frame set.

[0018] Optionally, step S14 comprises:

[0019] Step S141: slidingly pairing the view angle normalized image set in groups of two frames in time sequence, and performing pixel difference calculation on each frame pair of the sliding pairing result to obtain an inter-frame difference map;

[0020] Step S142: performing spatial connected domain extraction on the inter-frame difference map, eliminating regions with a connected domain area less than 20 pixels in the spatial connected domain, and constructing an inter-frame motion activation mask based on the maximum bounding box size and density distribution of the remaining regions;

[0021] Step S143: counting the proportion of active pixels in the inter-frame motion activation mask of each frame, and if the proportion of active pixels is less than a preset proportion threshold of 2%, marking the frame as a redundant frame and eliminating it; the remaining frames are retained as a candidate frame set;

[0022] Step S144: Set the time window as 2s, perform density distribution analysis on the frame time sequence in the candidate frame set, select the frame with the largest average motion activation area in the density distribution analysis result in each time window as the representative frame, and obtain a representative frame set;

[0023] Step S145: Perform confidence evaluation on the representative frame set, retain the frames with confidence evaluation results greater than 0.75, and obtain a high-confidence frame sequence.

[0024] Optionally, the step S2 of identifying the road boundary point set in the uniform acquisition frame set comprises:

[0025] Perform edge detection on the corresponding RGB visible light image, depth map, and infrared thermal image in each frame in the uniform acquisition frame set respectively, and perform weight fusion on the edge detection results to obtain a fused edge response map;

[0026] Extract a main line segment set in the fused edge response map, classify the main line segment set according to angle distribution, and select a continuous line segment in the horizontal direction and with a length exceeding 20% of the image width as a candidate road boundary line segment;

[0027] Calculate the boundary symmetry degree of a left-right region line segment pair in the candidate road boundary line segment, retain the line segment pair with a boundary symmetry degree greater than 0.85, and obtain a boundary line segment pair;

[0028] Reverse map the boundary line segment pair back to each perspective frame in the uniform acquisition frame set, perform boundary position re-projection error evaluation on the mapping result, retain the boundary line segment pair with a projection error less than 4 pixels in each perspective frame, and obtain a verified boundary line segment pair;

[0029] Perform curve fitting on the verified boundary line segment pair to obtain a fitted curve, and perform equidistant interpolation sampling on both sides of the fitted curve to obtain a road boundary point set.

[0030] Optionally, the step S2 of detecting dynamic targets and static targets in each time in the uniform acquisition frame set comprises:

[0031] Divide the road boundary point set into boundary points on both sides to obtain a left boundary point set and a right boundary point set;

[0032] Perform spline curve fitting on the left boundary point set and the right boundary point set respectively to obtain a left continuous boundary curve and a right continuous boundary curve;

[0033] Calculate a center line point at a corresponding position of each group of left continuous boundary curves and right continuous boundary curves, calculate a tangent vector and a normal vector of the center line point, and thereby construct a road coordinate system;

[0034] Projecting the spatial positions of the bounding boxes of each time in the unified collected frames to the corresponding road coordinate system to obtain target spatial position data to be analyzed;

[0035] Correlating the target spatial position data to be analyzed of each time across frames, marking and removing a target appearing less than 3 frames as a misjudgment target to obtain a target trajectory set;

[0036] Calculating the center point motion vector of each target in the target trajectory set between consecutive frames, and judging the motion state of the target according to the center point motion vector to obtain a motion state marked target set;

[0037] Overlapping the motion state marked target set with the road coordinate system for analysis, retaining only the targets in a boundary enclosing area, and dividing the retained targets into dynamic targets and static targets according to the motion state of the targets, the boundary enclosing area being an area enclosed by a left continuous boundary curve and a right continuous boundary curve.

[0038] Optionally, step S3 comprises:

[0039] Step S31: setting a sliding time window, segmenting the target trajectory corresponding to the dynamic target in the target trajectory set to obtain a short-time trajectory segment; calculating the motion vector sequence, speed change rate, acceleration feature and direction deviation angle of the short-time trajectory segment in a continuous time window to obtain a motion trajectory feature set;

[0040] Step S32: embedding each trajectory behavior in the motion trajectory feature set to obtain a trajectory behavior embedding vector set;

[0041] Step S33: removing an outlier trajectory behavior in the trajectory behavior embedding vector set, and marking a category label for the remaining trajectory behavior using a pre-defined target category dictionary to obtain the category of the dynamic target;

[0042] Step S34: annotating each target category in the category of the dynamic target, combining the time stamp and the spatial coordinates projected by the road coordinate system to perform spatial node coding to obtain a spatial node coding set;

[0043] Step S35: filling in the semantic continuity between frames on the time axis in the spatial node coding set, and constructing the motion boundary of each target on the spatial axis, thereby constructing a road semantic partition map;

[0044] Step S36: projecting the road semantic partition map into a standard grid map established by the road coordinate system, thereby constructing a road boundary semantic distribution map.

[0045] Optionally, step S4 comprises:

[0046] Step S41: Obtain the two-dimensional coordinate point of each semantic node in the road coordinate system according to the road boundary semantic distribution map, bind the image frame and timestamp in the high-confidence frame sequence closest to the two-dimensional coordinate point, and obtain a semantic image binding node set;

[0047] Step S42: Sort all nodes in the semantic image binding node set by timestamp, and aggregate them into frame segments according to a sliding time window to obtain a road boundary dynamic distribution map;

[0048] Step S43: Calculate the nearest neighbor spatial distance between dynamic targets and static targets in the road boundary dynamic distribution map, predict the nearest neighbor spatial distance trend of adjacent frames, construct an interaction graph between dynamic targets and static targets, and obtain a target interaction graph;

[0049] Step S44: Select the static target node affected by the dynamic target from the target interaction graph, perform a change trend analysis, and obtain a static target motion probability set;

[0050] Step S45: For a static target with a change probability higher than 0.7 in the static target motion probability set, perform a continuous time window confidence verification. If the change probability of two consecutive time windows is higher than 0.7, mark the static target as a motion trend target, and thus obtain the static target motion state change.

[0051] Optionally, the predicted target motion state overlap time point in step S5 includes:

[0052] According to the spatial threshold 1.5 m and the time overlap window 3 s, an interaction pair of a candidate dynamic target and a static target is constructed from the static target motion state change and the dynamic target of the road boundary dynamic distribution map, and a candidate target pair set is obtained;

[0053] Perform constant speed-mutation prediction on the static target in the candidate target pair set to obtain a static target state prediction trajectory set;

[0054] Perform multi-time step trajectory extrapolation on the dynamic target in the candidate target pair set to obtain a dynamic target state prediction trajectory set;

[0055] According to the static target state prediction trajectory set and the dynamic target state prediction trajectory set of each group, calculate the Euclidean distance between the two prediction trajectories at each time point, mark the time point set with all Euclidean distances less than 0.8 m, and obtain a prediction conflict time point set;

[0056] Perform motion overlap risk judgment on the prediction conflict time point set. If the prediction conflict time point between the two prediction trajectories in the group is more than 0.5 s, it is determined as a motion overlap risk time point, and thus a potential motion conflict time segment set is obtained;

[0057] The time point with the minimum spatial distance in the potential motion conflict time slice set is taken as the predicted overlap time point, and an overlap time point is obtained.

[0058] The application introduces structured semantic modeling, trajectory dynamic analysis, target interaction recognition and active prompt feedback in the security scene, breaks through the technical bottleneck of existing systems relying on image visual observation, response lag and passive interaction, and has significant improvement in display accuracy, judgment depth, response efficiency and intervention ability. In the initial stage of video processing, the motion activation mask is constructed through inter-frame pixel activation analysis, and 2% is set as the redundancy judgment threshold, which effectively filters out the background images with insufficient motion and significantly reduces the interference of redundant information, improving the system processing efficiency and resource scheduling ability. At the same time, the representative frame with the largest average activation area is selected within the 2s time window, and the confidence threshold of the reserved frame is set to 0.75, so that the display information is always concentrated in the most active and significant key moment, enhancing the focus of event perception and facilitating the operator to quickly locate the abnormal area in the large scene. In terms of spatial structure modeling, the fusion edge response graph is formed through multi-modal image fusion (RGB, infrared, depth) to avoid the interference problems of single image source such as light, shielding or texture; and the boundary symmetry (threshold set to 0.85) and the re-projection error (less than 4 pixels) are jointly constrained to ensure that the extraction of road boundary has geometric symmetry and multi-view consistency, which fundamentally improves the spatial expression ability of security scene. The subsequent semantic nodes constructed based on the road boundary and the binding with high-confidence frames enable the complex and unstructured dynamic scene to realize spatial semantic mapping and visual alignment. Compared with the image coordinate system or world coordinate system commonly used in existing security systems, the road coordinate system has obvious structural advantages and practical effect improvement. In the traditional security system, the image coordinate system only relies on pixel-level position for target labeling, lacks understanding of the real road structure, and makes it difficult to abstractly express the spatial relationship of the target, especially when the camera perspective changes or the region structure is complex, the target behavior path and traffic logic cannot be effectively modeled; while the world coordinate system can provide certain spatial positioning ability, but its construction depends on high-precision external parameter calibration and environment modeling, which is difficult to adapt, and lacks close association with road semantic structure, making it difficult to support dynamic target behavior constraint judgment and rule reasoning in actual traffic structure. The road coordinate system of the application automatically generates the centerline coordinate system consistent with the actual road direction and with physical semantic meaning through the steps of extracting road boundary point set, fitting continuous boundary curve, constructing centerline and direction vector, which not only has real-time construction ability, but also retains the road geometric properties and direction logic, and can be used as a unified spatial reference datum for position reduction, motion correlation and conflict analysis of multiple targets. Its structure not only has strong adaptability and high semantic correlation, but also can significantly enhance the intelligent level of security visualization system in target recognition, behavior prediction and event warning, breaking through the limitations of existing coordinate system in security complex scene in terms of expression and weak interaction logic. It can effectively make up for the shortcomings of existing security visualization means in spatial structure expression and target semantic organization.By constructing the left continuous boundary curve and the right continuous boundary curve based on the boundary points, and combining the tangent vector and normal vector information of the corresponding center line, a directional and hierarchical road coordinate reference system is formed, which can realize the spatial homing and semantic classification of the target in the monitoring picture. This coordinate system not only improves the structured degree of monitoring data, but also through the standardized mapping of the target position, makes the dynamic target and static obstacle can be intuitively displayed and dynamically tracked in a unified spatial framework, significantly reducing the cognitive burden of artificial identification of complex alarm conditions. In addition, with the help of the road coordinate system, the target behavior path can be predicted and conflict detected, which can identify potential traffic risks in advance and provide accurate security decision basis for high-density traffic areas or emergency passage scenes. In the aspect of behavior analysis, the extraction of short-time speed vector, direction angle and acceleration features in the dynamic target trajectory, by setting conditions such as speed mean>0.1m / s, standard deviation<0.05, etc., makes the system can identify objects with potential moving intention from low-speed or stationary targets. For trajectory prediction, the disturbance trajectory branch of static target combined with disturbance angle ±15° and speed disturbance ±25% is constructed, so that the prediction model can not only depict the regular stable behavior, but also include the feasibility of mutation behavior, greatly reducing the misjudgment rate of unpredictable factors. And for dynamic targets, by extracting global behavior and spatial neighbor joint features, a more reasonable and interactive-aware reasoning of future trajectory is achieved. The key lies in the construction of interaction atlas and the identification of motion overlap risk. By setting the spatial screening distance to 1.5m, the dangerous judgment distance to 0.8m and the minimum conflict time to 0.5s, the real conflict event between dynamic and static targets is identified, avoiding false alarm due to accidental approach, ensuring the practicality and accuracy of the prompt. On this basis, the static target with high mutation probability is verified for two time windows, only when the continuous stable trend appears, it will be marked as a trend target, which improves the noise resistance and behavior judgment stability of the system. Finally, through 5G communication, the conflict time point determined is transmitted to the target terminal, and the voice prompt module is linked to actively issue instructions such as "please be careful to avoid", compared with the traditional system relying only on monitoring images and background checking, the fusion of intelligent identification, target positioning and localized intervention is realized. Especially in the pedestrian-vehicle mixed scene, such intelligent terminals can be deployed on non-motor vehicles, smart bracelets or fixed facilities to realize real-time prompting and distributed security linkage at the target level. Its beneficial effects lie in not only enhancing the self-awareness and self-avoidance ability of the target itself, but also truly expanding the security equipment from "passive collection" to "active guidance", integrating the values of traffic safety judgment, spatial dynamic understanding and individual behavior warning.

[0059] Optionally, the application also provides an application for security visualization display system for executing the application for security visualization display method as described above, the application for security visualization display system comprising:

[0060] a data stream compression module, configured to acquire a security device collected data stream and compress the security device collected data stream to obtain a unified collected frame set;

[0061] a target classification module, configured to identify a road boundary point set in the unified collected frame set, and take the road boundary point set as a main axis of a road coordinate system to detect dynamic targets and static targets at each moment in the unified collected frame set;

[0062] a dynamic target category distinguishing module, configured to distinguish categories of the dynamic targets according to motion state changes of the dynamic targets at each moment in the unified collected frame set, and perform semantic labeling on the categories of the dynamic targets to obtain a road boundary semantic distribution graph;

[0063] a static target motion prediction module, configured to bind corresponding image slices and time stamps to each node in the road boundary semantic distribution graph to obtain a road boundary dynamic distribution graph, and analyze interaction conditions of the dynamic targets and the static targets according to the road boundary dynamic distribution graph to identify and predict motion state changes of the static targets;

[0064] a target collision early warning module, configured to predict an overlapping time point of target motion states based on the motion state changes of the static targets and the dynamic targets, and transmit the overlapping time point to a communication device of the static target through the Internet of Things to perform a voice prompt task.

[0065] The application is applied to a security visual display system, which can implement any one of the security visual display methods, is used as a medium for joint operation and signal transmission between modules, and completes the security visual display method, and the modules in the system cooperate with each other, so that the display precision and response efficiency of the security visual display are improved.

[0066] Optionally, the application further provides a computer readable storage medium storing a computer program, and the computer program is executed by a processor to implement the security visual display method. BRIEF DESCRIPTION OF DRAWINGS

[0067] Other features, objects and advantages of the application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:

[0068] Fig. 1 a step flowchart of the security visual display method of the application;

[0069] Fig. 2 a detailed step flowchart of step S1 in the application;

[0070] Fig. 3 a schematic diagram of a road coordinate system in step S2 in the application;

[0071] The objectives, features and advantages of the present application will be further illustrated with reference to the embodiments, with reference to the accompanying drawings. DETAILED DESCRIPTION

[0072] The technical method of the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0073] In addition, the accompanying drawings are only schematic illustrations of the present application, and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0074] It should be understood that although the terms "first", "second" and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be referred to as a second element, and similarly a second element can be referred to as a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0075] To achieve the above-mentioned object, please refer to Figs. 1 to 3 The present application provides a security and protection visual display method, which comprises the following steps:

[0076] Step S1: acquiring security and protection device collected data stream, and compressing the security and protection device collected data stream to obtain a unified collected frame set;

[0077] In this embodiment, 5 sets of security equipment are deployed in the main entrance area of a residential community, each set including an RGB high-definition camera, a TOF depth sensor, and an infrared thermal imaging unit, and data synchronization is performed through a unified edge collection node. During the collection process, 30 frames of RGB images, depth maps, and infrared images are synchronously acquired per second, while the device's attitude angle and lens parameters are recorded. First, the collected raw image stream is preliminarily synchronized through the camera's local timestamp, and the error tolerance range is set to ±20 ms. For image frames with time differences, spatial alignment is performed based on the attitude angle change trend, and missing frames are reconstructed through linear interpolation. After completing time synchronization, unified perspective conversion processing is performed using the camera's intrinsic matrix and device attitude to standardize all image frames to the front view. Subsequently, the normalized image frames are pairwise differenced in groups of two frames using a time sliding window approach, the brightness, depth, and thermal radiation changes of each pixel in adjacent frames are calculated, and static frames are removed based on the connected region area and the proportion of active pixels. In the remaining frames, a frame with the most active motion is selected every 2 seconds, and frames with a confidence score greater than 0.75 are retained through motion region confidence evaluation. Finally, the RGB images, depth maps, thermal maps, and their timestamps are spliced into a unified data frame structure to form a unified collection frame set.

[0078] Step S2: Identify the road boundary point set in the unified collection frame set, and use the road boundary point set as the main axis of the road coordinate system to detect dynamic targets and static targets at each time in the unified collection frame set.

[0079] In this embodiment, the obtained unified collection frame set is first processed by using an image processing method based on edge gradient to identify possible boundary line segments for each frame of image, wherein the RGB image uses Sobel edge extraction, the depth map uses edge gradient enhancement, and the infrared image uses heat intensity change rate mapping. The edge response maps of the three images are fused in a weight ratio of 3:2:1 to generate a fused edge map. In the map, a set of continuous line segments with a direction approximately horizontal and a length exceeding 20% of the image width is extracted, and the boundary symmetry degree of each pair of left and right line segments is calculated, and only the line segment combination with a symmetry degree greater than 0.85 is reserved. Subsequently, the line segment pairs are projected into multi-view images, the re-projection error is evaluated, and only the line segment pairs with an error less than 4 pixels in all view frames are reserved. The line segment pairs that pass the verification are subjected to quadratic curve fitting to obtain the boundary curve of the road on both sides, and the road region is uniformly interpolated based on the curve to generate a road boundary point set. When constructing the road coordinate system, the left and right boundary points are respectively subjected to cubic spline fitting, the tangent vector and normal vector of each center point are extracted, and a spatial reference system is established. The center point of the target detection box in each frame is projected into the road coordinate system, and the target center point sequence of all frames is associated, and the minimum number of consecutive frames is set to 3 frames, and only the continuously appearing target is reserved. The inter-frame motion vector size of these targets is calculated, and the target with a motion speed greater than 0.05 m / s is marked as a dynamic target, and the rest is a static target. Only the targets located in the fitted boundary enclosed region are reserved for classification.

[0080] Step S3: According to the motion state change of the dynamic target at each time in the unified collection frame set, the category of the dynamic target is distinguished, and the category of the dynamic target is semantically labeled, to obtain a road boundary semantic distribution map;

[0081] In this embodiment, in the unified frame set, for the trajectory data of the dynamic target determined, a sliding time window of 1.5 seconds is set to intercept the short-time trajectory segment of each dynamic target. For each trajectory, the motion direction change angle, speed increment and acceleration change amplitude between each frame are extracted, and the direction deviation rate is calculated to form a motion trajectory feature vector. The vector is mapped into an embedding vector structure with a dimension of 64, representing the differences in speed, direction and time continuity of different behavior patterns. After removing abnormal behavior trajectories using outlier detection, the remaining vectors are matched with a target behavior category dictionary constructed in advance, which includes “fast pedestrian”, “slow pedestrian”, “scooter”, “small electric vehicle” and other 8 types of target behavior. The determined category label is bound with its spatial projection coordinates and timestamp to form a spatial semantic node, and a frame filling operation is performed on the time axis to ensure semantic continuity. In the spatial dimension, the boundary envelope of the area where each type of target appears is extracted to establish the target behavior boundary. A semantic partition map is constructed in the entire road area and projected into the road coordinate grid to output the final road boundary semantic distribution map. Each grid unit in the distribution map represents the probability density of the appearance of a certain type of dynamic target and accurately corresponds to the road coordinates.

[0082] Step S4: binding the corresponding image slice and timestamp of each node in the road boundary semantic distribution map to obtain a road boundary dynamic distribution map; analyzing the interaction between dynamic targets and static targets according to the road boundary dynamic distribution map to identify and predict the motion state change of static targets;

[0083] In this embodiment, the two-dimensional coordinate points of each node in the obtained road boundary semantic distribution map are matched with the corresponding image frames in the high-confidence frame sequence to take the nearest timestamp image as the binding image of the semantic node. All binding nodes are arranged in ascending order of timestamp and aggregated into frame segments in a sliding window manner (2 seconds per window) to form a time-sequenced road boundary dynamic distribution map. In the distribution map, the position coordinates of dynamic targets and static targets are extracted respectively for each frame, and the spatial distance between the nearest neighbors is calculated. By comparing the spatial distance trend of adjacent frames, it is evaluated whether there is a close behavior. All target pairs with close trend are constructed as a group of interaction edges to form a target interaction graph. For the static target nodes in the interaction graph, the frequency and direction consistency of the interference by dynamic targets are counted, and the change probability of the future 3 seconds is calculated. The change probability threshold is set to 0.7, and if the probability of a static target exceeds the threshold in two consecutive sliding windows, it is considered to have a motion trend. Finally, the time period, spatial position and corresponding relationship with the interaction target of the batch of motion trend targets are output as the basis for subsequent conflict judgment.

[0084] Step S5: predicting the overlap time point of the target motion state based on the static target motion state change and the dynamic target motion state change, and transmitting the overlap time point to the communication device of the static target through the Internet of Things to perform a voice prompt task.

[0085] In this embodiment, after obtaining the motion trend of the static target, the motion trend is matched with the predicted trajectory of the dynamic target. The spatial overlap threshold is set to 1.5 meters, and the time window is set to 3 seconds. All target pairs that may interact within the time and space range are screened out. For each group of static targets, the central position trajectory of the last 4 seconds is extracted first, and the mean and standard deviation of the speed vector are calculated. If the mean speed is greater than 0.1 m / s and the standard deviation is less than 0.05, it is determined that it is a stable low-speed trend target, and a constant-speed linear trajectory is extrapolated. For targets with a mutation probability higher than 0.7, a trajectory branch containing ±15° angle and ±25% speed disturbance is constructed. In parallel, the dynamic target uses the direction, speed and acceleration features in its continuous trajectory, combined with the behavior features of the 5 spatial nearest neighbors in the trajectory neighborhood, to predict the future multi-time step trajectory. The Euclidean distance of the two types of trajectories is calculated every second within the prediction time window, and all time points below 0.8 meters are screened out to determine the potential conflict points. If the continuous conflict time period exceeds 0.5 seconds, the target group is marked as having a risk of motion overlap, and the time point with the smallest spatial distance in the conflict time period is selected as the final overlap time point. This time point information is transmitted to the embedded communication terminal or application software of the corresponding static target through the 5G Internet of Things connection, triggering the voice module to issue a "please be careful to avoid" voice prompt, and realizing active security prompt intervention. The static target can be a motor vehicle integrated with a communication terminal, a non-motor vehicle with a mobile intelligent terminal, or a living body carrying a networked smart bracelet with a voice prompt function. The design of sending a voice prompt to a static target is mainly based on the risk of accidental movement or conflict of a static target in the environment due to the influence of a dynamic target. Static targets usually include parked motor vehicles, parked non-motor vehicles, or personnel carrying smart bracelets. These objects are relatively fixed in the environment but have potential movement trends. By predicting the trajectory overlap time point of the static target and the dynamic target in real time, the warning information is transmitted to the communication device of the static target through the Internet of Things in a timely manner, which can realize early intervention and warning of potential dangers. This not only improves the initiative of safety protection and reduces the probability of collision or congestion events, but also effectively improves the response efficiency and intelligent management level of the security system, ensuring the safety and smoothness of personnel and vehicles in complex dynamic environments.

[0086] Optionally, step S1 comprises:

[0087] Step S11: Obtain the security device collection data stream and device auxiliary state data through the security device deployed in the security area, wherein the security device collection data stream includes an RGB visible light image sequence, a depth image sequence, and infrared thermal imaging data;

[0088] In this embodiment, four sets of multi-modal security monitoring devices are deployed in the main entrance area of a residential community, each set of device including a 200-megapixel color video camera (RGB visible light), a TOF depth perception camera with a resolution of 640x480, and an infrared thermal imager with a thermoelectric sensing array. These devices are arranged in a triangular staggered layout, covering the intersection area within a range of 3 to 12 meters, with an installation height of 3.2 meters and a downward inclination of the camera angle of 15°. Each set of device is configured with a local edge computing unit for collecting image sequences and auxiliary state data, including the collection timestamp of each frame (accuracy 0.001s), device number, current attitude angle (pitch angle, yaw angle, roll angle), and lens focal length, aperture value, and other internal parameter information. The collection frequency is uniformly set to 30 frames per second, and all collected data is transmitted in real time through the network port to the regional edge server for processing.

[0089] Step S12: Align the device timestamps of the security device collection data stream, and perform frame reconstruction interpolation on the timestamp alignment processing result according to the camera time sequence in the device auxiliary state data to obtain a synchronized and aligned image frame sequence;

[0090] In this embodiment, the image sequences uploaded by each device in the security device collection data stream are preliminarily time-aligned according to the collection timestamps, and the frame synchronization tolerance range is set to ±10ms. If the maximum time difference between different modal image frames exceeds this range, the frame is determined as an invalid frame and discarded directly. On this basis, according to the camera time sequence provided by the device, it is determined whether there is frame loss or time sequence gap in the frame sequence. For the time points that can be interpolated, according to the attitude difference and image brightness gradient of the previous and next two frames, a weighted reconstruction method is used to synthesize the interpolated frame, ensuring that the interpolated frame is consistent with the original frame in terms of spatial structure and light intensity distribution. All processed frames are renumbered with the same frame number and labeled with a uniform time tag, and finally form a 30-frame-per-second, three-channel synchronized image sequence, which serves as the basis for subsequent analysis.

[0091] Step S13: According to the pose parameters and lens internal parameters of each collection device in the device auxiliary state data, perform spatial view transformation processing on the synchronized and aligned image frame sequence to generate a view normalized image set;

[0092] In this embodiment, the obtained synchronized alignment image frame sequence is extracted to obtain the corresponding camera pose and lens parameters. According to the Euler angles and three-dimensional displacement values provided by the shooting device, the camera extrinsic matrix of the current frame is constructed, and the intrinsic matrix is constructed by combining the focal length, principal point position and radial distortion coefficient of the lens. Using this set of parameters, each image is back projected to a unified world coordinate reference plane, and is transformed to a set overhead view, the goal is to make the main viewing axis of all images perpendicular to the road surface, and the horizontal projection does not deviate. For infrared images and depth maps, linear interpolation is used to complete the hole region, and spatial resampling means is used to keep each modal image having the same resolution (unified to 720x1280 pixels) after transformation. The finally generated view normalized image set has spatial consistency and is suitable for subsequent cross-modal joint analysis.

[0093] Step S14: performing inter-frame difference feature extraction on the view normalized image set, removing redundant frames without target motion changes, and effectively sampling and compressing the remaining frames to generate a high-confidence frame sequence.

[0094] In this embodiment, the view normalized image set is processed frame by frame. First, the image frame change is analyzed in a time sliding window (window width 1 second, step length 0.5 second). In each time period, the difference operation is performed on the two consecutive image frames, the luminance gradient change of the RGB image, the geometric mutation area of the depth image, and the temperature hotspot moving track of the infrared image are calculated. Set the threshold: when the difference response value (normalized to take the maximum value of three channels) of an image frame is less than 0.05, and the continuous value below the threshold is more than 5 frames, the frame is marked as a frame without significant change and is removed. In the remaining image frames, the frame with the most active motion features in each sliding window is selected as the representative frame, and contour stability analysis is performed on the target area of the frame, only the frames with more than 80% of the motion area edge integrity and more than 400 pixels of the target area are reserved. The image frames reserved after this compression sampling constitute a high-confidence frame sequence, and the corresponding modal images are stored respectively.

[0095] Step S15: data splicing and packaging of the RGB visible light image, depth map, infrared thermal image and timestamp in the high-confidence frame sequence to obtain a unified acquisition frame set.

[0096] In this embodiment, the obtained high-confidence frame sequence is extracted frame by frame according to the timestamp, and the RGB image, depth map and infrared image are extracted, and the images are uniformly sized to 720x1280 pixels and kept in alignment order. Then, the three modal images in each frame are spliced and packaged according to the "RGB-DEP-IR-TS" format. Among them, the RGB image and the depth map are compressed and encoded with 8-bit grayscale, the infrared image maintains 16-bit thermal radiation accuracy, and the timestamp is attached to the header of the frame file as metadata. Finally, the uniform image frame structure unit is packaged, each unit occupies about 1.2MB of memory, has the characteristics of time consistency, uniformity of view angle and motion saliency. All frame structures are composed of uniform collection frame sets in time sequence, which are used as basic data input sources for subsequent boundary recognition and target tracking.

[0097] Optionally, step S14 comprises:

[0098] Step S141: slidingly pairing the view angle normalized image set in a group of two frames in time sequence, and performing pixel difference calculation on each frame pair of the sliding pairing result to obtain an inter-frame difference map;

[0099] In this embodiment, the view angle normalized image set is first strictly sorted according to the timestamp to ensure the continuity and time consistency of the image frame sequence. Then, a sliding window of two frames is used to traverse the entire sequence to form frame pairs. For each frame pair, first convert the RGB images of the two frames to grayscale images, extract the luminance channel through the weighted average formula (0.299xR+0.587xG+0.114xB) to weaken the influence of color difference. Then, the luminance difference of the corresponding pixel position is calculated to generate a difference image (size same as the original image, 720x1280 pixels), whose value range is 0 to 255, reflecting the luminance change at the pixel level. Set the difference threshold to 30, and the pixels whose difference value exceeds the threshold are marked as "motion pixels", and the rest are considered as static. The difference image is stored in 8-bit grayscale format to form an inter-frame difference map. The difference map is the basis for subsequent motion region extraction and reflects the regions that change significantly over time.

[0100] Step S142: performing spatial connected component extraction on the inter-frame difference map, removing regions with an area less than 20 pixels in the spatial connected component, and constructing an inter-frame motion activation mask based on the maximum bounding box size and density distribution of the remaining regions;

[0101] In this embodiment, the inter-frame difference map obtained in step S141 is subjected to spatial connectivity analysis to screen the motion region. Specifically, an 8-neighborhood-based connected region detection method is used to cluster all positions with pixel values greater than a threshold (30) to form a plurality of spatially connected regions. For each connected region, the total number of pixels and the size of the circumscribed bounding box are counted. Connected regions with an area less than 20 pixels are removed to avoid misjudgment caused by environmental noise, light fluctuations and device jitter. For the remaining connected regions, the width-height ratio of the maximum bounding box and the density distribution are further calculated, and the density is defined as the number of activated pixels in the connected region divided by the area of the bounding box. Using this information, an inter-frame motion activation mask is constructed, which is a binary image with a size of 720x1280, wherein the activated region pixels are marked as 1 and the non-activated region pixels are marked as 0. The mask accurately depicts the spatial range of the moving object in the image and excludes the interference region.

[0102] Step S143: The proportion of activated pixels in the inter-frame motion activation mask of each frame is counted. If the proportion of activated pixels is less than a preset proportion threshold of 2%, the frame is marked as a redundant frame and removed; the remaining frames are retained as a candidate frame set;

[0103] In this embodiment, the proportion of activated pixels in the motion activation mask corresponding to each frame is counted, and the calculation method is the number of activated pixels divided by the total number of pixels (921,600). When the proportion is less than the preset threshold of 2% (i.e., less than 18,432 activated pixels), it is determined that the frame lacks significant motion, and it is marked as a redundant frame and excluded from subsequent processing. For example, a frame of an empty road with no vehicle passing through is usually removed, greatly reducing the data volume. Frames with an activated pixel proportion ≥ 2% are retained to form a candidate frame set, providing candidate samples for subsequent key frame screening. This process ensures that processing resources are concentrated on dynamic significant frames, improving efficiency while reducing misjudgment.

[0104] Step S144: Set the time window to 2s, and perform density distribution analysis on the time sequence of frames in the candidate frame set. In each time window, select the frame with the largest average motion activation area in the density distribution analysis result as the representative frame to obtain a set of representative frames;

[0105] In this embodiment, the obtained candidate frame set is divided into a plurality of non-overlapping time windows according to the time stamp, and the window width is 2 seconds. In each time window, the motion activation area (i.e., the number of activated pixels in the mask) of all candidate frames is counted, and the average activation area is calculated. The frame with the largest activation area in the time window is selected as the representative frame, which represents the most dynamic information in the time period. For example, if a 2-second window contains 30 frames, the representative frame is the frame with the most obvious motion and the most complete target. If the number of frames in the window is less than 5, all frames are retained to avoid missing key dynamics. This method ensures that the representative frame can fully reflect the dynamic characteristics of the time segment through time density analysis, facilitating subsequent fine processing.

[0106] Step S145: Perform confidence evaluation on the representative frame set, retain frames with confidence evaluation results greater than 0.75 to obtain a high-confidence frame sequence.

[0107] In this embodiment, the obtained representative frame set is subjected to multi-dimensional confidence evaluation to comprehensively determine the quality and reliability of each frame. The evaluation content includes image sharpness (using edge intensity gradient mean, threshold set to 0.8), spatial coherence of motion region (shape similarity of continuous multiple frame activation masks ≥ 0.75), and spatial consistency of multi-modal data (position deviation between RGB, depth and infrared images does not exceed 3 pixels). The above indexes are fused to generate a comprehensive confidence score. Frames with a score greater than 0.75 are identified as high-confidence frames and included in the high-confidence frame sequence. This sequence guarantees excellent picture quality and sufficient dynamic information, providing a solid data foundation for target recognition and behavior analysis. The final sequence size is usually reduced by more than 70% compared to the original frame number, greatly improving the subsequent computing efficiency.

[0108] Optionally, the step S2 of identifying the road boundary point set in the uniformly collected frame set comprises:

[0109] Edge detection is performed on the RGB visible light image, depth map and infrared thermal image corresponding to each frame in the uniformly collected frame set, respectively, and the edge detection results are fused with weights to obtain a fused edge response map;

[0110] In this embodiment, for each frame image in the uniformly collected frame set, edge extraction processing is performed on the RGB visible light image, depth map and infrared thermal image, respectively. The edge extraction uses a detection method based on gradient amplitude and direction to ensure that the road edge features can be accurately captured. The RGB image is processed based on a luminance image composed of three channels, and the depth map and infrared image are subjected to edge detection according to the depth value gradient and temperature gradient, respectively. The edge results of each channel are fused with corresponding weights, with the weights set to RGB 0.5, depth map 0.3, and infrared thermal map 0.2, to adapt to the differences in edge performance under different light and environmental conditions. The fused edge response map is saved in a binary image form, highlighting the continuity and clarity of the road boundary.

[0111] The main line segment set in the fused edge response map is extracted, and the main line segment set is classified according to the angle distribution. Continuous line segments in the horizontal direction and with a length exceeding 20% of the image width are selected as candidate road boundary line segments.

[0112] In this embodiment, the main line segment set is extracted from the fusion edge response map. The main line segment is formed by aggregating continuous edge pixels, and the length unit of the line segment is pixel, and the standard image width is 1280 pixels. The angle distribution of all line segments relative to the horizontal axis is calculated, and the line segments are divided into different angle intervals using angle classification. The line segments with an angle deviation within ±10 degrees and a length exceeding 256 pixels (20% of the image width) are selected as candidate road boundary line segments. These line segments represent the possible road boundaries in the form of horizontal and continuous edges in the image. Through this step, noise line segments and non-road edges are effectively filtered out.

[0113] The boundary symmetry of the left and right area line segment pairs in the candidate road boundary line segments is calculated, and the line segment pairs with a boundary symmetry greater than 0.85 are retained to obtain the boundary line segment pairs;

[0114] In this embodiment, the left and right area line segment pairs of the candidate road boundary line segments are divided, and the boundary symmetry is calculated. The boundary symmetry is a similarity index of the corresponding line segment pairs in the horizontal direction and the vertical position. The matching degree of the length, position and direction vector of the left and right line segments is calculated. Only the line segment pairs with a boundary symmetry greater than 0.85 are retained to ensure that the boundary line pairs exhibit high symmetry and stability, thereby enhancing the reliability of the boundary line segments. For example, for a group of candidate line segment pairs, if the left line segment is 512 pixels long, the right corresponding line segment is 508 pixels long, and the position deviation is less than 10 pixels, then the symmetry is high and meets the retention condition.

[0115] The boundary line segment pairs are reversely mapped back to each perspective frame in the unified collection frame set, and the boundary position re-projection error evaluation is performed on the mapping result. The boundary line segment pairs with a projection error less than 4 pixels in each perspective frame are retained to obtain the verified boundary line segment pairs;

[0116] In this embodiment, the verified boundary line segment pairs are mapped back to each perspective image frame in the unified collection frame set. The mapping uses the pose parameters and lens intrinsics of the collection device in the device auxiliary state data to convert the boundary line segments to different perspective image coordinate systems through a three-dimensional space projection model. The re-projection error between the mapped line segments and the actual image edge position is calculated for each perspective frame. If the re-projection error in all perspective frames is less than 4 pixels, it is determined that the boundary line segment pair is verified, which ensures that the boundary line segment is consistent in multiple perspectives and enhances the spatial accuracy.

[0117] The verified boundary line segment pairs are subjected to curve fitting to obtain a fitted curve, and equidistant interpolation sampling is performed on both sides of the fitted curve to obtain a road boundary point set.

[0118] In this embodiment, the boundary line segments that pass the verification are curve-fitted in the two-dimensional image plane, a smoothing fitting method based on spline curve is adopted to process the local discontinuity and slight noise of the boundary points, and a continuous and smooth road boundary curve is obtained. Along both sides of the fitted curve, equidistant interpolation sampling is performed at a fixed interval of 0.2 meters to generate a high-density road boundary point set, which includes the two-dimensional coordinates of each sampling point and the corresponding time stamp. The point set provides an accurate basis for the subsequent establishment of the road coordinate system and target positioning, ensuring the spatial accuracy of positioning and analysis.

[0119] Fig. 3 FIG. 1 is a schematic diagram of the road coordinate system in step S2 of the present application; as shown in FIG. 1, the road coordinate system can be divided into a plurality of key regions, including but not limited to: a left boundary point set 101, a road coordinate system 102, and a right boundary point set 103. The road coordinate system is a continuous coordinate system set; the left and right boundary point sets are named by the relative spatial position of the electric set with respect to the road boundary line. Since the road conditions and viewing angles of each road are different, the naming of the boundary point set is not limited in the present application. Fig. 3

[0120] The left boundary point set 101 is generated by spline fitting of the left boundary point set, which can effectively describe the spatial structure profile of the left side of the road and provide a spatial basis for boundary constraint and target classification.

[0121] The road coordinate system 102 is the road centerline formed by extracting the center point and calculating the tangent vector and normal vector at the corresponding position of the left and right boundary curves. The centerline not only reflects the geometric principal axis direction of the road, but also serves as a reference for subsequent spatial position mapping and coordinate system establishment. The tangent vector (T) and normal vector (N) representing the road centerline are clearly marked on each center point. The tangent vector represents the main driving direction of the road, and the normal vector provides a directional reference for constructing a local rectangular coordinate system. Between each pair of continuous boundary curves, a standardized spatial coordinate system can be constructed based on the centerline, providing a structured geometric basis for subsequent spatial position projection, target state recognition, and interactive analysis.

[0122] The right boundary point set 103 is generated by fitting the right boundary point set, which together with the left boundary curve forms the physical boundary framework of the road space.

[0123] Optionally, the step S2 of detecting dynamic targets and static targets in the unified collection frame set includes:

[0124] The road boundary point set is divided into left and right boundary points to obtain the left and right boundary point sets;

[0125] ​In this embodiment, the road boundary point set is grouped according to its relative position in the road space, thereby forming a left boundary point set and a right boundary point set. This operation can be performed by combining the edge change area in the depth map and the reflection feature in the infrared image for spatial position screening, and the sampling interval of the boundary point set can be set to 0.2 meters to ensure that the accuracy of the boundary curve meets the spatial modeling requirements.

[0126] The left boundary point set and the right boundary point set are respectively fitted with a spline curve to obtain a left continuous boundary curve and a right continuous boundary curve.

[0127] The center line point is calculated at the corresponding position of each group of left continuous boundary curve and right continuous boundary curve, the tangent vector and normal vector of the center line point are calculated, and thus the road coordinate system is constructed.

[0128] In this embodiment, the left boundary point set and the right boundary point set are respectively fitted. Taking 20 continuous boundary points on each side as a group, a continuous boundary curve is generated using a cubic interpolation method, which is respectively denoted as a left continuous boundary curve and a right continuous boundary curve. This curve can better reflect the trend of the road structure and be used for subsequent center line and coordinate system construction. On this basis, for any group of left and right boundary curves at corresponding positions, the equidistant center point between them is calculated to construct a center line point list. At each center point, the tangent direction is further obtained through the difference between the front and rear points; the corresponding normal direction is obtained through the vector perpendicular to the tangent direction, thereby forming a road coordinate system with the center line as the main axis. The road coordinate system is a continuous set composed of a series of local coordinate sub-units, covering the entire monitoring road section. Fig. 3 In this embodiment, the road coordinate system is numbered as 102, the left boundary curve is numbered as 101, and the right boundary curve is numbered as 103.

[0129] The spatial positions of the boundary boxes in each time in the unified collection frame set are projected onto the corresponding road coordinate system to obtain the target spatial position data to be analyzed;

[0130] In this embodiment, all target detection boundary boxes in the unified collection frame set are projected onto the center line tangent plane according to the road coordinate system to obtain a standardized position representation of the target relative to the road space. This position representation includes not only the spatial projection point coordinates, but also the transverse distance relative to the road center line and the longitudinal trajectory segment number, thereby facilitating subsequent target association. The error limit in the projection operation is set to a maximum deviation of not more than 0.3 meters to ensure the stability and accuracy of the position mapping.

[0131] The target trajectory set is obtained by performing cross-frame association on the target spatial position data to be analyzed at each time, and marking and removing the targets that appear less than 3 frames as misjudgment targets.

[0132] In this embodiment, to exclude incidental misjudgment, the target recognition results in the continuous frames are subjected to cross-frame consistency judgment. With a three-frame sliding window, the target detection results near the same position are subjected to position matching. If a target appears less than twice in three frames, it is determined to be misrecognized and eliminated. The remaining target bounding boxes will form trajectory segments, constituting a target trajectory set. The time interval between trajectory points is set to 0.5 seconds, which can ensure tracking continuity and real-time performance.

[0133] The center point motion vectors of each target in the target trajectory set between the continuous frames are calculated, and the target motion state is judged according to the center point motion vectors to obtain a motion state marked target set.

[0134] The motion state marked target set is subjected to overlap analysis with the road coordinate system, only the targets in the boundary enclosing region are retained, and the retained targets are divided into dynamic targets and static targets according to the target motion state, and the boundary enclosing region is a region enclosed by the left continuous boundary curve and the right continuous boundary curve.

[0135] In this embodiment, the center point offset vector between adjacent continuous frames of each target trajectory is extracted as the motion feature of the target. The target with a motion distance greater than 0.2 meters / second is marked as a dynamic state target, and the target with a motion distance less than 0.2 meters / second is processed as a static state target. The target motion state marking result is mapped to the road coordinate system again, and overlap analysis is performed with the closed area surrounded by the left boundary curve and the right boundary curve. The overlap analysis refers to the containment judgment of the spatial position (in the form of a center point) of each target in the road coordinate system with the road boundary area composed of the left fitting curve and the right fitting curve. The road boundary area is a closed belt-shaped polygon constructed according to the lateral projection width of the two side boundaries based on the center line. The method of judging whether the target is in the area includes: extracting the lateral distance coordinate d and the longitudinal line segment number s of the target in the road coordinate system, and judging whether d is located between the left / right boundary (the boundary point distance is calculated as 0.2 meters, and the error tolerance range of the belt width is ±0.1 meters). Only the targets that meet the condition that the center point falls within the area are retained for subsequent interactive analysis. Only the target trajectories within the area surrounded by the boundary curve are retained, and the dynamic state targets are divided into dynamic targets, and the static state targets are divided into static targets for subsequent analysis. The area is the closed road segment space composed of numbers 101 and 103 in the figure. The targets outside the area may be background noise or off-site targets, which need to be eliminated. The purpose of such processing is to exclude false detection results caused by background noise, reflection interference or environmental changes in the external area of the road, and to ensure that the analysis targets are concentrated on functional roads such as fire passages. Through the target set after the boundary overlap constraint, the accuracy of dynamic / static target classification can be effectively improved, and reliable spatial structure support can be provided for subsequent operations such as constructing interactive pairs and predicting overlap time.

[0136] Optionally, step S3 comprises:

[0137] Step S31: setting a sliding time window, segmenting the target trajectories corresponding to the dynamic targets in the target trajectory set to obtain short-time trajectory segments; calculating the motion vector sequence, speed change rate, acceleration feature and direction offset angle of the short-time trajectory segments in the continuous time window to obtain the motion trajectory feature set;

[0138] In this embodiment, the trajectory data of the target trajectory concentrated dynamic target is sorted according to the timestamp, the sliding time window length is set to 5 seconds, the window step is 1 second, the trajectory data is divided into multiple short-time trajectory segments by using continuous time periods. For each short-time trajectory segment, the motion vector of each frame center point is calculated, which is the difference vector of the position coordinates of the adjacent two frames, with the unit of meter / second. At the same time, the velocity change rate and the acceleration feature are calculated based on the motion vector, the velocity change rate is the relative change percentage of the continuous frame velocity, and the acceleration is calculated by the difference of the velocity change rate. In addition, the motion direction offset angle of each frame is calculated, which quantifies the small change of the target moving direction, with the unit of degree. The above motion parameters constitute the comprehensive motion trajectory feature set of the short-time trajectory segment, and the data format is a time sequence matrix, which contains four-dimensional information of position, velocity, acceleration and direction.

[0139] Step S32: embedding each trajectory behavior in the motion trajectory feature set to obtain a trajectory behavior embedding vector set;

[0140] In this embodiment, each short-time trajectory behavior in the above motion trajectory feature set is mapped to a fixed-dimensional behavior description space. The space uses a pre-trained trajectory behavior representation model, which is based on a multi-layer time series data coding unit and can convert multi-dimensional time series input into an embedding vector with a length of 128. The input includes the above motion feature matrix, which is encoded frame by frame and globally pooled to output the corresponding trajectory behavior embedding vector. The embedding vector set contains the behavior description of all short-time trajectories and is stored as an N×128 matrix, where N is the number of trajectory segments. This representation helps to uniformly describe motion behaviors of different lengths and complexities.

[0141] Step S33: removing outlier trajectory behaviors in the trajectory behavior embedding vector set, and labeling the remaining trajectory behaviors with class labels using a pre-defined target class dictionary to obtain the class of the dynamic target;

[0142] In this embodiment, outlier behavior screening is performed on the trajectory behavior embedding vector set. Threshold screening based on vector similarity distribution is used to remove abnormal embeddings with a cosine similarity lower than 0.4 to the majority of vectors, and the removed objects usually correspond to abnormal or noisy trajectory segments. The remaining embedding vectors are matched according to a pre-defined target class dictionary, which includes categories such as "pedestrian", "bicycle", "motor vehicle" and their standard embedding samples. Through nearest neighbor matching, the trajectory behavior is assigned a corresponding class label. The class label sequence of the dynamic target generated by this classification process assists subsequent semantic labeling and analysis, ensuring accurate and realistic traffic scene classification.

[0143] Step S34: labeling each target class in the class of the dynamic target, and combining the timestamp, the spatial coordinates projected in the road coordinate system to perform spatial node coding to obtain a spatial node coding set;

[0144] In this embodiment, after the dynamic target category labeling is completed, spatial node codes are generated for each target trajectory in combination with the trajectory timestamp and the space coordinates projected into the road coordinate system. The coding format is a four-tuple (target ID, timestamp, X coordinate, Y coordinate), where X and Y are two-dimensional coordinates in the road coordinate system, with units of meters. The spatial node code set is stored in table form, supporting spatiotemporal queries and indexing. This data structure lays the foundation for subsequent construction of a spatiotemporal semantic graph, ensuring accurate correspondence of spatiotemporal information of dynamic targets and supporting continuous semantic inference.

[0145] Step S35: Inter-frame semantic continuous filling is performed on the time axis in the spatial node code set, and motion boundaries of various targets are constructed on the spatial axis, thereby constructing a road semantic partition graph.

[0146] Step S36: The road semantic partition graph is projected into a standard grid graph established in the road coordinate system, thereby constructing a road boundary semantic distribution graph.

[0147] In this embodiment, for the spatial node code set, first, semantic continuous filling is performed on the frames with missing data in the time axis direction. Linear interpolation is used to complete the spatial coordinates and category labels of the missing time points, avoiding semantic discontinuity. In the spatial axis direction, motion boundaries are constructed according to the distribution of targets of the same category, and target motion range boundaries are formed by calculating the minimum envelope curve between spatial nodes. The semantic relationship between time and space is integrated to form a multi-level road semantic partition graph, which is refined to a 0.5-meter grid size, and the hierarchical structure reflects the category and motion state of dynamic targets. Finally, the semantic partition graph is mapped to a standard two-dimensional grid graph (size set to 500x500 pixels) established in the road coordinate system, achieving spatial visualization distribution of semantic information, and obtaining a road boundary semantic distribution graph, which provides a spatial semantic basis for subsequent target interaction analysis.

[0148] Optionally, step S4 comprises:

[0149] Step S41: According to the road boundary semantic distribution graph, the two-dimensional coordinate points of each semantic node in the road coordinate system are obtained, the image frames and timestamps in the closest high-confidence frame sequence of the two-dimensional coordinate points are bound, and a semantic image binding node set is obtained.

[0150] In this embodiment, the two-dimensional coordinate points of all semantic nodes are extracted from the road boundary semantic distribution map, and the coordinate points are based on the pre-established road coordinate system with a unit of meters. For each two-dimensional coordinate point, the high-confidence frame sequence is traversed to find the image frame closest to the node in terms of time stamp and spatial position, and the Euclidean distance is used as the spatial matching standard, and the time difference is limited within ±0.5 seconds. After successful matching, the node is bound to the corresponding image frame and time stamp to generate a semantic image binding node set. The set structure is a multi-field table containing node ID, two-dimensional coordinate, image frame index and time stamp, facilitating subsequent spatio-temporal correlation analysis and query.

[0151] Step S42: sort all nodes in the semantic image binding node set by time stamp, and aggregate them into frame segments by a sliding time window to obtain a road boundary dynamic distribution map;

[0152] In this embodiment, all nodes in the semantic image binding node set are first sorted in ascending order according to the time stamp to ensure the time sequence continuity. Then, the sliding time window length is set to 3 seconds and the step length is set to 1 second, and the nodes are divided into multiple frame segments according to the time window. Each frame segment contains all node data within the time window. This frame segmentation processing can effectively aggregate node information similar in time and space to generate a road boundary dynamic distribution map. The map takes frame segments as the time unit and the spatial distribution is based on the node position, supporting subsequent dynamic target time sequence analysis. When constructing the road boundary dynamic distribution map, in addition to aggregating the semantic nodes of the dynamic target according to the time window, the spatial position points of the static target in the same time period are also added to the map as independent nodes. These static target nodes are also bound to the corresponding time stamp and two-dimensional coordinate to form a node set parallel to the dynamic nodes. This ensures that the dynamic distribution map becomes a unified spatio-temporal representation structure containing two types of targets.

[0153] Step S43: calculate the nearest neighbor spatial distance between the dynamic target and the static target in the road boundary dynamic distribution map, predict the trend of the nearest neighbor spatial distance in the adjacent frame, and construct an interaction graph between the dynamic and static targets to obtain a target interaction graph;

[0154] In this embodiment, the nearest neighbor spatial distance between the dynamic target and the static target nodes in the road boundary dynamic distribution map is calculated. The distance calculation uses the Euclidean distance of two-dimensional coordinates as the measurement, and the neighbors within 2 meters are considered as potential interaction objects. Through the spatial distance trend analysis of the continuous time frame segments, the relative motion trend between the targets is determined by the distance change rate between the previous and next two frames. Based on these data, an interaction graph between the dynamic and static targets is constructed, and the graph structure is a weighted network of edges between nodes, and the weight reflects the distance and motion trend between the targets, forming a comprehensive target interaction graph.

[0155] Step S44: Selecting static target nodes affected by dynamic targets from the target interaction graph, performing variable trend analysis to obtain a set of static target motion probabilities;

[0156] In this embodiment, static target nodes affected by dynamic targets are selected from the interaction graph, and variable trend analysis is performed on these nodes. The variable trend first collects the continuous spatial position change data of the target nodes within the sliding time window, calculates the distance change rate and direction consistency of the target and the surrounding dynamic targets, and combines the target's own speed and acceleration characteristics to comprehensively evaluate whether the target has a trend from static to mobile. The analysis uses a weighted fusion method, taking distance shortening trend, speed increasing trend and motion direction stability as main indicators, and reduces noise interference through time series smoothing processing, finally outputting a probability value reflecting the target motion possibility. In this evaluation process, the sliding time window length is 4 seconds, the state change frequency and trend of the target within the time window are calculated to ensure the timeliness and accuracy of the probability value.

[0157] Step S45: For static targets in the static target motion probability set with a variable probability higher than 0.7, perform continuous time window confidence verification. If the variable probability of two consecutive time windows is higher than 0.7, mark the static target as a motion trend target, thereby obtaining the static target motion state change situation.

[0158] In this embodiment, static targets with a motion probability higher than 0.7 are further subjected to continuous time window confidence verification. Both consecutive time windows are required to meet the probability threshold to ensure the continuity and stability of the variable prediction. Through this double window verification mechanism, false positives caused by incidental noise are prevented. Static targets that meet the conditions are marked as motion trend targets, forming the final static target motion state change situation output. This state labeling result is used to guide subsequent dynamic interaction early warning and active prompting functions, ensuring accurate perception of potential mobile static targets by the security system.

[0159] Optionally, the predicted target motion state overlap time point in step S5 includes:

[0160] According to the spatial threshold 1.5m and the time overlap window 3s, the candidate dynamic target and static target interaction pairs are constructed from the static target motion state change situation and the dynamic target of the road boundary dynamic distribution graph, and a candidate target pair set is obtained;

[0161] In this embodiment, the spatial threshold is set to 1.5 meters, and the time overlap window length is 3 seconds. By screening the motion state change of the static target and the spatio-temporal position of the dynamic target in the road boundary dynamic distribution map, the dynamic-static target pair that may interact is matched. The specific operation includes traversing the two-dimensional coordinates and motion state information of the static target in the specified time period, combining the spatial position of the dynamic target in the same time window, and constructing a candidate target pair set. The set contains all target pairs that meet the spatial distance and time overlap conditions, ensuring the spatio-temporal consistency and accuracy of the interaction analysis.

[0162] Perform constant speed-mutation prediction on the static target in the candidate target pair set to obtain a static target state prediction trajectory set;

[0163] In this embodiment, for the static target in the candidate target pair, the constant speed and mutation trend combination motion state prediction is performed according to its latest motion trajectory data. The prediction process includes extracting the mean speed and standard deviation in the continuous time window from the static target center position sequence to determine the motion stability; under the premise of meeting the stable low-speed motion, linear extension is made along the speed direction, and based on the motion mutation probability, candidate trajectory branches with ±15 degree direction disturbance and ±25% speed change are generated to ensure that the prediction results cover possible motion changes. The trajectory data structure is indexed by time stamp and position coordinates.

[0164] Perform multi-time step trajectory extrapolation on the dynamic target in the candidate target pair set to obtain a dynamic target state prediction trajectory set;

[0165] In this embodiment, for the dynamic target in the candidate target pair, the multi-step time point trajectory extension prediction is performed by using its trajectory data sequence, combining the speed, acceleration and direction change characteristics in the continuous time window, and cooperating with the motion state of the most 5 spatial nearest neighbors in the neighborhood. The multi-time step prediction output includes a time series vector of the future position of the dynamic target, and the data structure is in the form of time series, supporting time-by-time position query and update, and ensuring fine-grained reflection of the motion trend of the dynamic target.

[0166] According to the static target state prediction trajectory set and the dynamic target state prediction trajectory set of each group, the Euclidean distance between the two prediction trajectories at each time is calculated, and all time point sets with Euclidean distance less than 0.8m are marked to obtain a prediction conflict time point set;

[0167] In this embodiment, within the prediction trajectory time window of each target pair, for each second sampling time point, the two-dimensional Euclidean distance between the static target prediction trajectory point and the dynamic target prediction trajectory point is calculated, and the time points with a distance less than 0.8 meters are collected as potential prediction conflict time points. The distance threshold is set according to the actual target size and safety distance to ensure effective identification of risk moments of spatial proximity. All time points that meet the conditions are centrally managed to form a set of prediction conflict time points.

[0168] The set of prediction conflict time points is subjected to motion overlap risk determination. If the prediction conflict time points between the two prediction trajectories in the set exceed 0.5 seconds, it is determined to be a motion overlap risk time point, and a set of potential motion conflict time segments is obtained.

[0169] In this embodiment, the continuous time period in the set of prediction conflict time points is evaluated. If the prediction conflict points in any continuous time period last for more than 0.5 seconds, the time period is determined to be a motion overlap risk time period, and a set of potential motion conflict time segments is constructed. Time duration determination helps to filter occasional short-distance proximity events, enhancing the stability and reliability of the determination.

[0170] The time point with the smallest spatial distance in the set of potential motion conflict time segments is taken as the prediction overlap time point, and the overlap time point is obtained.

[0171] In this embodiment, the time point with the smallest spatial distance in the set of potential motion conflict time segments is selected as the predicted target motion state overlap time point. This time point is considered as the most critical conflict moment, which is used for subsequent safety warning and intervention operations. The overlap time point is sent in real time to the communication device of the static target through the Internet of Things communication channel, ensuring timely voice prompts, assisting the target in safe avoidance, and improving the proactive protection capability of the overall security system.

[0172] Especially important is that the constant speed-mutation prediction specifically includes:

[0173] From the static target in the set of candidate target pairs, the center point position sequence in the continuous time window is extracted.

[0174] In this embodiment, for each static target in each set of candidate target pairs, based on the standard two-dimensional plane in the road coordinate system, the high-confidence frame sequence in the unified collection frame set is used to extract the center position point sequence of the static target within the past 5 seconds (indexed by timestamp) with a sampling interval of 0.5 seconds, thereby forming a continuous trajectory segment. The position point is converted into ground two-dimensional coordinates (in meters) through the center of the spatial detection box in the RGB image combined with the depth map or laser radar mapping calibration, ensuring the geometric accuracy and comparability of the trajectory.

[0175] The average speed and the speed standard deviation of the speed vector between two continuous time windows in the center point position sequence are calculated. If the average speed is greater than 0.1 m / s and the speed standard deviation is less than 0.05, it is considered that the static target between the two time windows has a stable low-speed trend, and the time window pair with the stable low-speed trend is extracted, so as to obtain the target speed vector feature;

[0176] In this embodiment, the center point position sequence is divided into a plurality of adjacent time window segments (the window length is 1 second), for example, 1-2 seconds, 2-3 seconds, 3-4 seconds, etc. The speed vector from the previous window to the next window is calculated for each pair of time windows. The speed vector is obtained by dividing the coordinate difference between the terminal point and the starting point in the window by the window interval time (i.e. 1 second), and each segment of the speed vector is recorded in turn. For all continuous speed vector sequences, the average (representing the trend speed) and the standard deviation (representing the stability) of the speed module length of each segment are calculated. When the average speed of a certain trajectory segment is greater than 0.1 m / s and the speed standard deviation is less than 0.05 m / s, it is considered that the target continuously moves at a low speed and uniform speed trend in this time period, and it is determined as a stable low-speed trend segment.

[0177] Based on the average speed of the target speed vector feature and the current position in the candidate target pair set, a constant speed linear extrapolation is performed to obtain a constant speed prediction trajectory.

[0178] In this embodiment, the current frame is taken as the reference time point, and the last position point and the speed vector at the tail of the nearest stable trend segment are selected as the initial conditions for constant speed linear extrapolation. When performing constant speed linear extrapolation, the end position of the static target in the stable low-speed motion trend segment is taken as the starting point, and the average speed vector calculated in this time period is combined to calculate the future motion path of the target at a fixed time interval. The prediction length is usually set to 3 seconds, and the time interval is 0.5 seconds, that is, 6 future time position points will be calculated in turn. These position points form a straight line trajectory based on the assumption that the current speed and direction do not change. While generating the trajectory, it is also checked whether each prediction point is in the passable area defined in the road semantic map to avoid falling into non-passable areas such as isolation belts and fence areas. If the trajectory meets all the passability requirements, the constant speed trajectory is considered as an effective prediction path, and is used as the basis for subsequent risk assessment and disturbance trajectory construction. The prediction period is fixed at 3 seconds, the sampling interval is still 0.5 seconds, and a prediction position sequence of 6 time points is generated, which is called a constant speed prediction trajectory.

[0179] The mutation probability of the constant speed prediction trajectory is evaluated, and disturbance trajectory branches are generated by integrating the direction disturbance angle ± 15° and the speed disturbance amplitude ± 25% based on the constant speed prediction trajectory.

[0180] In this embodiment, the behavior characteristics of static targets are used, including the acceleration change amplitude, direction fluctuation frequency, trajectory curvature mutation times in the past 5 seconds, and other behavior factors, to evaluate the mutation probability through a weighted decision rule. The mutation probability is a continuous value (0-1), and a high value indicates that the static target is more likely to have a dramatic change in future movement direction or speed. In the disturbance construction, disturbance parameters are introduced in the speed and direction dimensions respectively: based on the current average speed, a disturbance amplitude of ±25% is set, i.e., the speed is modified to 0.75 times and 1.25 times the original speed; in the movement direction, ±15 degrees of deflection is added to the original speed vector direction to form multiple direction disturbance paths. Thus, 9 disturbance branch trajectories (including 1 main path and 8 disturbance paths) can be combined, each trajectory still has 6 nodes formed by 0.5 second sampling to ensure time consistency. All disturbance trajectories have a path type label field, such as "main trajectory", "positive deflection + fast", "negative deflection - slow", etc., to facilitate path source discrimination in subsequent conflict analysis.

[0181] If the mutation probability is <0.7, output the linear trajectory in the constant speed prediction trajectory as the static target state prediction trajectory set;

[0182] If the mutation probability is ≥0.7, output the disturbance trajectory branch labeled with the main branch as the static target state prediction trajectory set.

[0183] In this embodiment, if the mutation probability result is less than 0.7, it is considered that the target behavior is stable, and the constant speed prediction trajectory described above can be directly used as the final prediction trajectory set; but if the mutation probability is not less than 0.7, it is considered that the target may have uncertain behavior, and trajectory disturbance needs to be introduced. The output path set will be selected according to the mutation probability judgment: if stable, output 1 linear trajectory in the constant speed prediction trajectory; if unstable, output the complete main branch path set. This prediction trajectory set will be used for spatial proximity analysis at each time point with the trajectory set of the dynamic target to identify potential conflict intervals and support accurate calculation of motion state overlap time points. The whole process is composed of technically stable, structurally clear, and scene-adaptive static target trajectory extrapolation process through parameter-controllable disturbance modeling, clear speed / direction rules, and mutation condition judgment threshold.

[0184] Especially important is that the multi-time step trajectory extrapolation is specifically:

[0185] The spatial center point sequence is extracted from the dynamic target trajectory set, and the speed vector, direction angle, and acceleration information of the continuous time window are extracted to obtain the trajectory feature vector sequence;

[0186] In this embodiment, the continuous spatial center point sequence of each target in the current time window is extracted from the generated dynamic target trajectory set. The corresponding time interval of the sequence is usually set to 0.5 seconds, and the total length is 4 seconds, thereby forming a short trajectory segment with a length of 8 frames. For each frame, the corresponding velocity vector, motion direction angle (based on the north direction, ranging from 0-360°) and linear acceleration value calculated from the previous and next frames are synchronously extracted, and finally an information sequence containing time sequence, i.e. a trajectory feature vector sequence, is formed, with a dimension of NxD, N being the number of frames and D being the motion feature dimension of each frame.

[0187] The trajectory feature vector sequence is subjected to deep time sequence coding to output a dynamic target global time sequence feature vector.

[0188] In this embodiment, after the trajectory feature vector sequence is constructed, a set of time sequence feature extraction units is used to code the time-related information. The extraction process is realized through a specific structure, such as a gating structure or a nested convolution channel, to extract continuousness, change rate and acceleration period features that can reflect the overall motion trend of the target. The finally extracted time sequence coding result is output in the form of a fixed-length vector as the global time sequence feature vector of the target at the current stage.

[0189] The spatial neighbor set of the dynamic target in the dynamic target global time sequence feature vector at the current time in the dynamic target trajectory set is extracted, and the trajectory feature vector of the spatial neighbor set is extracted to obtain a neighbor trajectory feature vector.

[0190] In this embodiment, at the current time, the spatial neighbor set of the target is retrieved from the dynamic target set. The selection principle of the neighbor set is to perform spatial search in the road coordinate system with a radius of 2 meters, and at most 5 effective neighbor targets are retained. For these neighbor targets, the trajectory feature vector sequence in the current time window is also extracted, and is sorted from near to far according to the spatial distance. The sorted neighbor trajectory feature vector is spliced and fused with the global time sequence feature vector of the target itself to form a joint feature description for coding the behavior background of the target in the local environment.

[0191] The dynamic target global time sequence feature vector and the neighbor trajectory feature vector are spliced, and the position is gradually predicted for the future time window to obtain a dynamic target state prediction trajectory set.

[0192] In this embodiment, according to the future prediction time window (usually 2 seconds, a total of 4 prediction time points, every 0.5 seconds) set according to the current scene, a step-by-step prediction operation is performed. The process inputs the joint features and the historical motion state into a time sequence position updating module, calculates the target future position point by point in time, and ensures that each step of the prediction result has continuity and physical rationality. The final output of a set of continuous two-dimensional coordinate points constitutes the motion trajectory of the dynamic target in the prediction time window, that is, the dynamic target state prediction trajectory set. The trajectory set not only reflects the potential movement trend of the target, but also contains its adaptive response in the local crowd dynamic environment, providing accurate input for subsequent spatial overlap analysis.

[0193] Optionally, the present application also provides an application for a security and protection visual display system for executing the application for a security and protection visual display method as described above, the application for a security and protection visual display system comprising:

[0194] a data stream compression module for acquiring a security and protection device collected data stream and compressing the security and protection device collected data stream to obtain a unified collection frame set;

[0195] a target classification module for identifying a road boundary point set in the unified collection frame set and taking the road boundary point set as a road coordinate system main axis to detect dynamic targets and static targets at each moment in the unified collection frame set;

[0196] a dynamic target category distinguishing module for distinguishing the category of the dynamic target according to the motion state change of the dynamic target at each moment in the unified collection frame set and performing semantic labeling on the category of the dynamic target to obtain a road boundary semantic distribution graph;

[0197] a static target motion prediction module for binding corresponding image slices and time stamps of each node in the road boundary semantic distribution graph to obtain a road boundary dynamic distribution graph, analyzing the interaction between the dynamic target and the static target according to the road boundary dynamic distribution graph, and identifying and predicting the motion state change of the static target;

[0198] a target collision warning module for predicting an overlap time point of the target motion state based on the motion state change of the static target and the motion state change of the dynamic target, and transmitting the overlap time point to the communication device of the static target through the Internet of Things to perform a voice prompt task.

[0199] Optionally, the present application also provides a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the application for a security and protection visual display method as described above.

[0200] Therefore, the embodiments should be regarded, at any point, as being exemplary and not limiting, the scope of the application being defined by the appended claims and not by the above description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.

[0201] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for visualizing security applications, characterized in that, Includes the following steps: Step S1: Acquire the data stream collected by the security device and compress the data stream to obtain a unified collection frame set; Step S2: Identify the road boundary point set in the unified acquisition frame set, and use the road boundary point set as the principal axis of the road coordinate system to detect dynamic and static targets at each moment in the unified acquisition frame set; Step S3: Based on the changes in the motion state of dynamic targets at various times in the unified collection of frames, distinguish the categories of dynamic targets, and perform semantic annotation on the categories of dynamic targets to obtain the semantic distribution map of road boundaries; Step S4: Bind the corresponding image slices and timestamps to each node in the road boundary semantic distribution map to obtain the road boundary dynamic distribution map; analyze the interaction between dynamic and static targets based on the road boundary dynamic distribution map, and identify and predict the changes in the motion state of static targets; Step S5: Based on the changes in the motion state of the static target and the changes in the motion state of the dynamic target, predict the overlap time point of the target's motion state, and transmit the overlap time point to the communication device of the static target through the Internet of Things to perform the voice prompt task.

2. The security visualization display method according to claim 1, characterized in that, Step S1 includes: Step S11: Obtain the security equipment data stream and auxiliary status data through the security equipment deployed in the security area, wherein the security equipment data stream includes RGB visible light image sequence, depth map sequence and infrared thermal imaging data; Step S12: Perform device timestamp alignment processing on the data stream collected by the security equipment, and perform frame reconstruction interpolation on the timestamp alignment processing result based on the camera time series in the device auxiliary status data to obtain a synchronized image frame sequence; Step S13: Based on the pose parameters of each acquisition device and the lens intrinsic parameters in the auxiliary status data of the equipment, perform spatial viewpoint transformation processing on the synchronized image frame sequence to generate a viewpoint normalized image set. Step S14: Perform inter-frame difference feature extraction on the viewpoint normalized image set, remove redundant frames without target motion changes, and effectively sample and compress the remaining frames to generate a high-confidence frame sequence. Step S15: The RGB visible light image, depth map, infrared thermal image and timestamp in the high confidence frame sequence are stitched and encapsulated to obtain a unified acquisition frame set.

3. The security visualization display method according to claim 2, characterized in that, Step S14 includes: Step S141: Perform two-frame sliding pairing of the viewpoint normalized image set in chronological order, and perform pixel difference calculation on each frame pair of the sliding pairing result to obtain the inter-frame difference map. Step S142: Perform spatial connected component extraction on the inter-frame difference map, remove regions with a connected component area of ​​less than 20 pixels, and construct an inter-frame motion activation mask based on the maximum bounding box size and density distribution of the remaining regions. Step S143: Calculate the proportion of activated pixels in the inter-frame motion activation mask of each frame. If the proportion of activated pixels is less than 2% of the preset proportion threshold, the frame is marked as a redundant frame and discarded; the remaining frames are retained as a candidate frame set. Step S144: Set the time window to 2s, perform density distribution analysis on the frame time series in the candidate frame set, and select the frame with the largest average motion activation area from the density distribution analysis results in each time window as the representative frame to obtain the representative frame set. Step S145: Perform a confidence assessment on the representative frame set, and retain the frames with a confidence assessment result greater than 0.75 to obtain a high-confidence frame sequence.

4. The security visualization display method according to claim 1, characterized in that, Step S2 involves identifying the set of road boundary points within the unified acquisition frame set, including: Edge detection is performed on the RGB visible light image, depth map, and infrared thermal image corresponding to each frame in the unified acquisition frame set, and the edge detection results are weighted and fused to obtain a fused edge response map. Extract the set of main line segments from the fused edge response map, classify the set of main line segments according to their angular distribution, and select continuous line segments in the horizontal direction with a length exceeding 20% ​​of the image width as candidate road boundary line segments; Calculate the boundary symmetry of the left and right region line segment pairs in the candidate road boundary line segments, and retain the line segment pairs with a boundary symmetry greater than 0.85 to obtain the boundary line segment pairs; The boundary line segment pairs are reverse-mapped back to each view frame in the unified acquisition frame set. The boundary position reprojection error is evaluated on the mapping result. Boundary line segment pairs with projection errors less than 4 pixels in each view frame are retained to obtain the verified boundary line segment pairs. The verification boundary line segments are fitted with curves to obtain the fitted curve; equidistant interpolation sampling is performed on both sides of the fitted curve to obtain the road boundary point set.

5. The security visualization display method according to claim 1, characterized in that, Step S2 involves detecting dynamic and static targets at various times within the unified acquisition frame set, including: The road boundary point set is divided into two boundary points, resulting in the left boundary point set and the right boundary point set; Spline curve fitting is performed on the left boundary point set and the right boundary point set respectively to obtain the left continuous boundary curve and the right continuous boundary curve; For each set of continuous boundary curves on the left and right sides, calculate the centerline points at the corresponding positions, and calculate the tangent vector and normal vector of the centerline points to construct the road coordinate system. The spatial positions of the bounding boxes at each time point in the unified collection of frames are projected onto the corresponding road coordinate system to obtain the spatial position data of the target to be analyzed. Cross-frame correlation is performed on the spatial location data of the target to be analyzed at each time point. Targets that appear in less than 3 frames are marked as misjudged targets and removed to obtain the target trajectory set. Calculate the center point motion vector of each target in the target trajectory set between consecutive frames, and determine the target motion state based on the center point motion vector to obtain the motion state labeled target set; The motion state marker target set is overlaid with the road coordinate system, and only targets within the boundary area are retained. The retained targets are divided into dynamic targets and static targets according to their motion state. The boundary area is the region jointly enclosed by the left continuous boundary curve and the right continuous boundary curve.

6. The security visualization display method according to claim 1, characterized in that, Step S3 includes: Step S31: Set a sliding time window, segment the target trajectory corresponding to the dynamic target in the target trajectory set to obtain a short-time trajectory segment; calculate the motion vector sequence, velocity change rate, acceleration characteristics and direction offset angle of the short-time trajectory segment within the continuous time window to obtain the motion trajectory feature set; Step S32: Embed each trajectory behavior in the motion trajectory feature set to obtain a trajectory behavior embedding vector set; Step S33: Remove outlier trajectory behaviors from the trajectory behavior embedding vector set, and use a predefined target category dictionary to label the remaining trajectory behaviors with category labels to obtain the category of the dynamic target; Step S34: Label each target category in the category of dynamic targets, and combine it with its timestamp and spatial coordinates projected in the road coordinate system to perform spatial node encoding, thereby obtaining a spatial node encoding set; Step S35: Perform inter-frame semantic continuity filling on the time axis of the spatial node encoding set, and at the same time construct the motion boundaries of various targets on the spatial axis, thereby constructing a road semantic partition map; Step S36: Project the road semantic partition map onto a standard raster map established in the road coordinate system to construct a road boundary semantic distribution map.

7. The security visualization display method according to claim 1, characterized in that, Step S4 includes: Step S41: Obtain the two-dimensional coordinates of each semantic node in the road coordinate system based on the semantic distribution map of the road boundary, and bind the image frame and timestamp in the high-confidence frame sequence that is closest to the two-dimensional coordinate point to obtain the semantic image binding node set; Step S42: Sort all nodes in the semantic image binding node set by timestamp and aggregate them into frame segments by sliding time window to obtain a dynamic distribution map of road boundaries; Step S43: Calculate the nearest neighbor spatial distance between dynamic and static targets in the dynamic distribution map of road boundaries, predict the nearest neighbor spatial distance trend of adjacent frames, construct the interaction map between dynamic and static targets, and obtain the target interaction map. Step S44: Select static target nodes affected by dynamic targets from the target interaction graph, perform change tendency analysis, and obtain the static target motion probability set; Step S45: For static targets whose motion probability concentration changes more than 0.7, perform continuous time window confidence verification. If the change probability is higher than 0.7 for two consecutive time windows, then mark the static target as a motion trend target, thereby obtaining the change of the static target's motion state.

8. The security visualization display method according to claim 1, characterized in that, Step S5 includes predicting the overlapping time points of the target motion state: Based on a spatial threshold of 1.5m and a temporal overlap window of 3s, interactive pairs of candidate dynamic targets and static targets are constructed from the changes in the motion state of static targets and the dynamic distribution map of road boundaries, resulting in a set of candidate target pairs. Perform constant-rate-mutation prediction on the static targets in the candidate target pair set to obtain the static target state prediction trajectory set; Multi-time-step trajectory extrapolation is performed on the dynamic targets in the candidate target set to obtain the dynamic target state prediction trajectory set; Based on the static target state prediction trajectory set and the dynamic target state prediction trajectory set of each group, calculate the Euclidean distance between the two prediction trajectories at each moment, mark the set of all time points where the Euclidean distance is less than 0.8m, and obtain the set of prediction conflict time points. The motion overlap risk is determined for the set of predicted conflict time points. If the predicted conflict time point between two predicted trajectories in this set exceeds 0.5s, it is determined to be a motion overlap risk time point, and the potential motion conflict time segment set is obtained. The time point with the smallest spatial distance in the set of potential motion conflict time segments is taken as the predicted overlapping time point to obtain the overlapping time point.

9. A security visualization display system, characterized in that, For performing the security visualization display method as described in claim 1, the security visualization display system includes: The data stream compression module is used to acquire and compress the data stream collected by the security equipment to obtain a unified set of acquisition frames. The target classification module is used to identify the road boundary point set in the unified acquisition frame set, and to detect dynamic and static targets at each moment in the unified acquisition frame set, using the road boundary point set as the main axis of the road coordinate system. The dynamic target category differentiation module is used to differentiate the categories of dynamic targets based on the changes in the motion state of dynamic targets at various times in the unified collection of frames, and to perform semantic annotation on the categories of dynamic targets to obtain a semantic distribution map of road boundaries. The static target motion prediction module is used to bind corresponding image slices and timestamps to each node in the road boundary semantic distribution map to obtain the road boundary dynamic distribution map; based on the road boundary dynamic distribution map, it analyzes the interaction between dynamic and static targets, and identifies and predicts the changes in the motion state of static targets. The target collision warning module is used to predict the overlap time point of the target motion state based on the changes in the motion state of static targets and the changes in the motion state of dynamic targets, and transmits the overlap time point to the communication device of the static target through the Internet of Things to perform voice prompt tasks.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the security visualization display method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for predicting motion trail of non-motor vehicle and related equipment

    CN116224317A

  • Motion data analysis method and system based on image processing

    CN120411172A

  • Method for motion compensation in digital dynamic video

    RU2552139C1

  • Image processing method and apparatus

    US20210327074A1

  • Transitory salient attention capture to draw attention to digital document parts

    US20220284071A1

Cited By

  • Highway vehicle parking detection method and device based on unmanned aerial vehicle cruise

    CN121505493A

  • VR interaction data management method based on artificial intelligence

    CN121664203A

  • Abnormal dynamic detection method in feed preparation process

    CN121982616A

  • Method for detecting abnormal dynamics in a feed preparation process

    CN121982616B