Personnel unsafe behavior online identification method based on space-time Transform

By using cross-camera topology maps and GraphTransformer in collaboration with ReID for consistent decoding, the problem of misjudgment in cross-camera identity tracking and behavior recognition within the power plant campus was solved, enabling real-time and accurate identification and linkage of unsafe behaviors.

CN121305682APending Publication Date: 2026-01-09GUONENG CHONGQING WANZHOU ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511637536.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve real-time identity tracking and unsafe behavior identification across cameras within power plant parks, especially in complex scenarios where misjudgments are prone to occur. Furthermore, traditional methods suffer from computational delays and information loss.

Method used

An online identification method for unsafe human behavior based on spatiotemporal Transformer is adopted. By constructing a trajectory map and injecting appearance, time and topology information through cross-camera topology mapping, GraphTransformer and ReID collaborative consistency decoding, cross-camera identity tracking and behavior recognition can be achieved.

Benefits of technology

It enables real-time identification and immediate linkage of unsafe behaviors of personnel within the power plant park, significantly reduces the probability of cross-camera identity matching errors, reduces computational latency, and outputs structured global unsafe behavior event records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305682A_ABST
    Figure CN121305682A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power plant automatic management, and particularly relates to a space-time Transform-based personnel unsafe behavior online identification method, which comprises the following steps of: 1, setting a scene and solidifying a cross-camera topological graph, including establishing a node record and a directed edge record, accessing a video stream to generate a pedestrian bounding box set, a safety helmet state category and an appearance description vector, and forming a detection set of an identification period; step 2, an online reasoning process of topological gating three-channel GraphTransform (GraphTransform) and ReID (Reference Identifier) collaborative consistency decoding is executed; and step 3, online publishing and linkage: outputting a structured result after each identification period is finished, and sending the structured result to a scheduling end through a message bus to trigger a linkage strategy. According to the invention, cross-camera identity consistency identification and behavior continuity analysis can be realized in a complex industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power plant automatic management technology, specifically relating to an online identification method for unsafe human behavior based on spatiotemporal Transformer. Background Technology

[0002] With the continuous development of industrial production safety supervision technology, high-risk locations such as power plants, chemical plants, and metallurgical plants have widely deployed video surveillance systems for personnel safety status detection. However, existing technologies for identifying unsafe behaviors mainly rely on image analysis algorithms within the range of a single camera. These algorithms are typically based on convolutional neural networks for pose estimation or on behavior recognition models for video segment classification. They can identify local behaviors such as falling, running, and climbing within a single field of view, but they struggle to continuously track the same person across multiple cameras or identify persistent violations over a large area. For example, common single-camera detection systems can only determine whether someone is wearing a safety helmet within the current frame, but cannot determine whether that person continues to walk bareheaded through multiple areas after entering the plant area. This limits the effectiveness of the system in real power plant scenarios.

[0003] Another existing technology introduces pedestrian re-identification (ReID) methods, which achieve identity matching across multiple cameras by extracting pedestrian appearance features. Although ReID technology can identify the same person in cross-camera scenarios, it mostly relies on appearance similarity as the primary criterion, lacking constraints on temporal order and spatial topology. In power plant parks, there are strict path restrictions between different passages, corridors, and access control points. Relying solely on appearance similarity can easily lead to incorrect matches, especially when work clothes are uniform and environmental backgrounds are similar. Some systems have attempted to introduce temporal constraints, such as setting a maximum time interval across cameras, but this linear threshold constraint cannot adapt to different path lengths and different area traffic characteristics, still easily causing misjudgments. Furthermore, traditional ReID models generally operate independently of the behavior recognition module, unable to simultaneously identify the security behavior status of the same person while identifying their identity. This forces the algorithm to perform multiple inferences in real-time scenarios, increasing system latency and computational costs. Summary of the Invention

[0004] Therefore, the main objective of this invention is to provide an online identification method for unsafe personnel behavior based on spatiotemporal Transformer. Through joint optimization of cross-camera topology graphs, GraphTransformer, and ReID collaborative consistency decoding, it achieves real-time identity tracking and unsafe behavior identification of personnel across camera positions and areas within a power plant campus. This method first models the field of view of each camera as nodes, establishing directed edges based on actual traversable paths to form a cross-camera topology graph. Then, through trajectory graph construction, three-channel feature injection, and topology-gated GraphTransformer encoding, a stable joint context representation is generated by combining appearance, time, and topology information. Next, identity consistency decoding and event template matching mechanisms are used to synchronously output cross-camera IDs and global unsafe behavior events. Finally, a message bus is used to achieve online publishing of structured results and linkage with the scheduling end.

[0005] The technical solution adopted in this invention is as follows: A spatiotemporal Transformer-based online method for identifying unsafe human behaviors, comprising: Step 1: Scene setting and cross-camera topology map solidification, including establishing node records and directed edge records, and connecting video streams to generate pedestrian bounding box sets, safety helmet status categories and appearance description vectors, forming a detection set for a recognition cycle; Step 2: Execute the online inference process of topology-gated three-channel GraphTransformer and ReID collaborative consistency decoding, including: Step 2.1: Construct the trajectory map; Step 2.2: Perform three-channel feature injection; Step 2.3: Perform topology-gated three-channel GraphTransformer encoding; Step 2.4: Construct an identity consistency decoder; Step 2.5: Perform event decoder matching with template; Step 2.6: Generate collaborative and consistent joint optimization decisions; Step 3: Online publishing and linkage, including outputting structured results after each recognition cycle and sending them to the scheduling end via message bus to trigger linkage strategies.

[0006] Furthermore, the establishment of node records and directed edge records in step 1 includes: establishing a node record for each camera position. This node record contains the camera position number, the two-dimensional ground coordinate point, the number of vertices of the field of view polygon outline (no less than 8), and one entry direction code and one exit direction code. The value sets for the entry direction code and the exit direction code are as follows: One-to-one correspondence; a directed edge record is established based on the connectivity between the park's fixed access channels, access gates, and fence gates. This directed edge record includes the starting point gate number, the ending point gate number, the estimated travel time (lower limit 2 seconds and upper limit 60 seconds), and the access area code sequence.

[0007] Furthermore, the access video stream in step 1 generates a set of pedestrian bounding boxes, a helmet state category, and an appearance description vector, forming a detection set for a recognition cycle. This includes: accessing a 1920x1080 resolution video stream at 25 frames per second; executing a human bounding box generator on each frame to output a set of pedestrian bounding boxes, and performing non-maximum suppression based on an overlap area ratio of 0.50 or higher; executing a helmet state generator to output a helmet state category, where the set of values ​​for the helmet state category is... ; Execute a ReID encoder to output an appearance description vector with a length of 256; Collect the human body frame, safety helmet status category and appearance description vector of each camera position in a recognition cycle of 1 second, and form the detection set of this recognition cycle.

[0008] Further, step 2.1: Constructing the trajectory map includes: within each recognition cycle, the detections of the same target from the same camera position in adjacent frames are aggregated into groups of 3 frames. When the overlap area ratio between any two bounding boxes in the 3 frames reaches 0.70 or higher, a detection segment is formed. A trajectory map is built using the detection segments as nodes, and three types of directed edges are established: Type A: Time-adjacent edges from the same camera position, established with a time interval limit of 1 second and a node center point displacement limit of 100 pixels; Type B: Candidate edges across camera positions, established with the criterion that there is only a corresponding directed edge from the starting camera position to the ending camera position in the cross-camera topology map and the time difference between the two detection segments is within the expected travel time interval of the corresponding directed edge; Type C: Jump edges from the same camera position, established with a time interval limit of 3 seconds and a node center point displacement limit of 150 pixels.

[0009] Further, step 2.2: Performing three-channel feature injection includes injecting three sets of descriptions for each node and edge of the trajectory graph: appearance description, which is the normalized sequence of the 256-dimensional appearance description vector from the ReID encoder; time description, which is the start time stamp, end time stamp, and duration of the detected segment, and the duration is divided into intervals. Perform discrete coding; topological description, obtain discrete codes for gate number, cross-gate directed edge identifier, lower and upper limits of expected travel time, and travel area coding sequence.

[0010] Further, step 2.3: Performing topology-gated three-channel GraphTransformer encoding includes: sequentially stacking attention encoders for the appearance, time, and topology channels on the trajectory graph. Each channel contains two encoder layers, and each layer sequentially performs multi-head attention, feedforward update, and layer normalization. The appearance channel aggregates appearance descriptions on directed edges of type A and type C to form a local appearance context. The time channel performs message propagation on directed edges of type A, type B, and type C with time sequence as the key to form a time context. The topology channel introduces a topology gate, which performs three checks for each directed edge of type B: existence and cross-camera topology. Figure 1 If all three conditions are met—directed connectivity, the time difference between two detection segments being within the expected traversal time interval of the corresponding directed edge, and the encoding sequence of the traversing area being consistent—then the directed edge of type B is marked as a high-priority edge, and the directed edge of type B that is not marked as a high-priority edge is marked as a low-priority edge; after the three channels are executed sequentially, a joint context representation is generated for each node.

[0011] Further, step 2.4: Constructing an identity consistency decoder includes: performing set-growth identity grouping with the joint context representation as input. The process is as follows: an identity group is created using the node with the highest score as the anchor point, and new nodes are absorbed according to the following three criteria to form cross-camera IDs: Criterion 1 is that the appearance similarity score reaches 0.82 or above; Criterion 2 is that the time order is increasing and the time interval between the same camera position does not exceed 1 second, and the time interval between cross-camera positions is located within the expected traversal time interval of the corresponding directed edge; Criterion 3 is that there is a directed path from the previous node position to the candidate node position on the cross-camera topology graph and the number of path hops does not exceed 8; when multiple candidate nodes meet all three criteria at the same time, the candidate node with the earlier timestamp is selected, followed by the candidate node with the higher appearance similarity score; after the identity group is absorbed within this recognition period, a time-sorted camera position sequence and time sequence are formed.

[0012] Further, step 2.5: Executing the event decoder and template matching includes: for each completed identity group, performing event matching within the identification period based on the camera position sequence, time sequence, and helmet status category sequence. The event template contains three subcategories with clearly defined judgment conditions using numerical thresholds: Template A is bare-head through patrol, the judgment condition is that the identity group's helmet status category is bare-head at the first camera position and subsequently completes at least 4 camera position changes, forming a sequence of continuously passing through 5 camera positions with a cumulative dwell time of 20 seconds or more; Template B is high-voltage area timeout. The criteria for determining a stay are: the identity group stays continuously for 20 seconds or more within the set of camera positions marked with a high-voltage zone. The high-voltage zone marker is represented by the camera position attribute field isHV being 1. The Cth template is reverse passage, which is determined by: the identity group completing one camera position switch in the opposite direction of a directed edge in the cross-camera topology map, and the two ends of the camera positions are consistent with the start and end points of the corresponding directed edge. The event decoder generates an event record for the identity group that satisfies any template. The event record includes the event type code, the associated cross-camera ID, the camera position sequence, and the start and end timestamps.

[0013] Further, step 2.6: generating collaborative consistency joint optimization decisions includes: performing two rounds of collaborative consistency calibration within the same forward link; in round 1, the cross-camera ID and camera position sequence produced by the identity consistency decoder are used as the main output for initial event records and triggering a path completer. The path completer inserts virtual type B directed edges for identity groups with camera position sequence gaps and cross-camera topology connectivity within a time interval of no more than 15 seconds and returns to execute steps 2.3 and 2.4 for one incremental update; in round 2, the event records produced by the event decoder are used as the main output for camera position sequence calibration of identity groups triggering templates A and C, so that camera position switching corresponds one-to-one with the directed edges in the cross-camera topology; when the average edge score of all identity groups in this identification period reaches 0.80 or above or two rounds of calibration have been completed, the inference of this identification period ends and a cross-camera ID and a global list of insecure behavior events are generated simultaneously.

[0014] Furthermore, the online publishing and linkage in step 3 includes: outputting a structured result after each recognition cycle, which includes cross-camera ID strings, camera position sequence, entry and exit times of each camera position, a global list of unsafe behavior events, and event type codes; sending it to the dispatch terminal via the message bus; the dispatch terminal triggers the linkage strategy according to the event type code, pushes alarms to the gatekeeper and adjacent security terminals for the record corresponding to template A of the event type code, and calls the playback interface to play back the data in the camera position sequence order; the end-to-end processing delay target is 1 second or less.

[0015] By adopting the above technical solutions, this invention achieves the following beneficial effects: Through joint optimization of cross-camera topology graphs, GraphTransformer, and ReID collaborative consistency decoding, online identification and real-time linkage of unsafe behaviors of personnel within power plant parks are realized, outperforming existing technologies in overall architecture, information fusion method, and output mechanism. First, based on cross-camera topology graphs, this method explicitly models the paths between the fields of view of each camera as directed edges, constraining the identification results to occur only in physically reachable camera combinations within the algorithm, thus significantly reducing the probability of cross-camera identity matching errors. Second, the three-channel encoding mechanism of GraphTransformer jointly models appearance description, temporal sequence, and topological reachability, enabling the identification process to move beyond relying on a single visual similarity judgment and instead perform global reasoning based on spatiotemporal logical consistency, accurately tracking the continuous behavior of the same person in complex scenarios. Through topology gating, the model can automatically suppress feature propagation along unreasonable paths during operation, ensuring that the reasoning results remain coherent and interpretable. Furthermore, the introduction of ReID collaborative consistency decoding enables identity recognition and behavior recognition to be completed synchronously within the same network, avoiding the delays and information loss caused by traditional multi-stage processing, thus achieving second-level response within the recognition cycle. This method can not only output cross-camera IDs but also generate structured global records of insecure behavior events, including camera position sequences and start and end timestamps, facilitating rapid location of violation paths by the dispatching end. Finally, through the online publishing and linkage mechanism of the message bus, the recognition results can be pushed to gate and security terminals in real time, enabling automatic alarms and rapid response. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the cross-camera topology structure provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the principle of three-factor constraint determination and high / low priority edge selection for a topology gate provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the candidate head region for helmet detection provided in an embodiment of the present invention. Detailed Implementation

[0017] All features disclosed in this specification, or all steps in all disclosed methods or processes, may be combined in any way, except for mutually exclusive features and / or steps.

[0018] Any feature disclosed in this specification (including any appended claims and abstract) may be replaced by other equivalent or similar features, unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is merely one example of a series of equivalent or similar features.

[0019] A spatiotemporal Transformer-based online method for identifying unsafe human behaviors, comprising: Step 1: Scene setting and cross-camera topology map solidification, including establishing node records and directed edge records, and connecting video streams to generate pedestrian bounding box sets, safety helmet status categories and appearance description vectors, forming a detection set for a recognition cycle.

[0020] In one implementation, node records and directed edge records are first established within the power plant area. A node record is created for each camera position, containing the camera position number, a two-dimensional ground coordinate point, a field-of-view polygon outline, an entry direction code, and an exit direction code. The camera position number uses a combination of a fixed prefix and an incrementing number, such as "CAM-0001". The two-dimensional ground coordinate point is represented in meters on the unified reference plane of the park, measured using a rangefinder and total station at the centerline of the passageway; this facilitates direct conversion between subsequent path length and estimated travel time, reducing distortion caused by pixel scale. The field-of-view polygon outline uses no fewer than 8 vertices, drawn in the ground coordinate system based on the camera position, lens focal length, and obstruction boundaries. The reason for using 8 or more vertices is that the pipe corridors, trestle bridges, and railings in the power plant area form a polygonal passable area; a polygon with a low number of vertices would overextend or shrink the actual visible area, leading to false openings or closings in subsequent topology accessibility determinations. There is one entry direction code and one exit direction code. The set of direction code values ​​is {1,2,3,4}, which corresponds one-to-one with {North, East, South, West}. The direction codes are set according to the convention of north being up and east being right in the park's overall plan. Through the unified direction codes, the orientation of the machine positions in different construction areas can be mapped to a consistent global reference direction, which facilitates the splicing of cross-area paths.

[0021] The establishment of directed edge records is based on the connectivity between fixed access channels, access gates, and fence gates within the park. Each directed edge record includes the starting station number, the ending station number, the lower and upper limits of the estimated travel time, and the access area code sequence. The lower limit of the estimated travel time is set to 2 seconds, and the upper limit is set to 60 seconds; fine-tuning based on the actual length of each channel and ground conditions is allowed. The lower limit of 2 seconds is set to exclude obviously impossible human teleportation, thereby suppressing cross-station misconnections caused by image jitter or false detections; the upper limit of 60 seconds is set because cross-area movement within the factory area is affected by access control verification and turning waiting, and an excessively low upper limit would cut off the actual path. The access area code adopts the form of "R" prefix followed by 3 digits, such as R001, R002, R003. The access area code sequence is listed in the order of the channel from the starting station to the ending station, used to confirm whether the path crosses the same group of management areas during subsequent topology reachability comparisons. The advantage of using the pass area coding sequence is that it can separate the constraints of "visible" and "passable": even if the fields of view of two camera positions overlap, as long as the pass area codes are inconsistent, they will not be misjudged as passable paths.

[0022] After completing the above recording, the cross-camera topology map is loaded into the online inference service. To ensure time consistency across all camera positions, a network time protocol or a precision time protocol is used to synchronize the clocks of all camera positions, with timestamp errors controlled within ±20 milliseconds. The purpose of this is to allow direct comparison of detections from different camera positions within the same recognition cycle, thereby reducing the uncertainty of cross-camera time sequence from the source.

[0023] When inputting video streams, a resolution of 1920×1080 and a frame rate of 25 frames per second are used, with all image frames accompanied by timestamps accurate to milliseconds. The reason for choosing 25 frames per second is that in the power plant's traffic environment, the actual speed of pedestrians is typically between 1 and 2 meters per second. 25 frames per second allows the displacement of the center point between adjacent frames to remain on the order of tens to hundreds of pixels, supporting stable temporal aggregation without introducing unnecessary computational load. Each image frame is fed into a human bounding box generator, a safety helmet state generator, and a ReID encoder, sequentially obtaining a set of pedestrian bounding boxes, a safety helmet state category, and an appearance description vector.

[0024] The specific implementation process of the human bounding box generator is as follows. First, the input frame is scaled by its long side to maintain 1920 pixels, and color normalization is performed to reduce contrast differences under different lighting conditions. Image denoising filtering is then used to suppress high-frequency noise. Next, dense candidate regions are generated and human bodies are identified across the entire image. The candidate regions cover the area from the knees to the top of the head to ensure sufficient tolerance for accessories such as backpacks and handheld objects. After obtaining the candidate boxes, non-maximum suppression is performed, with the overlap area ratio threshold set to 0.50, retaining the candidate boxes with the highest confidence. The reason for setting the overlap area ratio threshold to 0.50 is that pedestrians in factory areas typically pass side-by-side with shoulder-width spacing. If the threshold is set too high, candidate boxes of different people will be incorrectly merged; if the threshold is set too low, too many duplicate boxes will be retained, leading to an overexpansion of the subsequent trajectory map. After non-maximum suppression, the set of pedestrian bounding boxes for the current frame is obtained.

[0025] The safety helmet state generator takes each pedestrian bounding box as input and first determines the head candidate region within the pedestrian bounding box. The head candidate region adopts a strategy of selecting a height ratio range from the top boundary downwards, preferably between 0.35 and 0.55 of the height from the top. This is because pedestrian posture changes in the image are mainly torso swings, while the head's vertical position within the bounding box is relatively stable. Selecting this range can simultaneously cover head-down, head-up, and slight side tilts, effectively avoiding interference from arms and tools. Edge contour aggregation and color saliency extraction are performed within the candidate region, prioritizing the detection of high-saturation, arc-shaped closed areas. Common colors of safety helmets in factory areas are white, yellow, red, and blue. These colors have saturation peaks and brightness distributions in the hue-saturation-brightness color space that are significantly different from skin and hair colors. Based on this, candidate regions are sorted by color saliency and further filtered using arc closure: when the edge closure ratio of a candidate region is not less than 0.70 and the color difference threshold reaches 30 (in the range of integers from 0 to 255), it is considered compliant with helmet wearing regulations; when no region meeting the above conditions is found within the candidate region, it is considered bare-headed. The advantage of this sequential filtering strategy is that even in the event of glare or a bright background, the dual constraints of color and shape can still exclude reflective areas, thus stably generating helmet status categories. The set of values ​​for helmet status categories is {bare-headed, compliant with helmet wearing}.

[0026] The appearance description vector is generated based on the pedestrian bounding box, which is then clipped and normalized. The pedestrian bounding box is extended by 10% on each side to preserve the texture information of clothing edges and carrying devices. The clipped image is then scaled to a fixed size of 256×128, and brightness equalization is applied to reduce the difference between backlighting and front lighting. A 256-bit appearance description vector is obtained via a ReID encoder, and this vector is normalized to ensure that descriptions from different camera positions and under different lighting conditions are of a uniform dimension. The 256-bit appearance description vector strikes a balance between computational resources and discriminative power: in scenarios like power plants where work clothes and protective clothing are predominantly worn, textures and color blocks are relatively regular. If the appearance description vector is too short, it will lead to insufficient cross-camera discriminative power; if the appearance description vector is too long, it will significantly increase the storage and comparison costs of the trajectory map in each recognition cycle.

[0027] To form a detection set for a recognition cycle, the system uses 1 second as one recognition cycle. Within that second, the pedestrian bounding box set, safety helmet status category, and appearance description vector for each camera position are written to the cache in timestamp order, and aggregated using the camera position number and timestamp as keys. The direct benefit of using 1 second as the recognition cycle is that a fixed window can be used to align the temporal relationships across camera positions: at the common passage speed in power plants, cross-camera events within 1 second rarely occur, thus reducing incomplete splicing caused by "frame interpolation"; at the same time, a 1-second cycle corresponds to 25 frames, ensuring that the pedestrian bounding box set within the same recognition cycle has sufficient temporal density, facilitating subsequent aggregation of detection segments in groups of 3 frames. At the end of the recognition cycle, the detection set for that recognition cycle is output, along with the recognition cycle number and start and end timestamps, for use in subsequent step 2.

[0028] In another optional implementation, the number of vertices in the field-of-view polygon contour is set to 12 or 16 to handle boiler rooms and substation courtyards with large-scale occlusion. Increasing the number of vertices improves the polygon's fit to obstacles and restricted areas, resulting in more precise accessibility determination across the camera topology map and correspondingly reducing the error of edges marked as low-priority in the topology gating process in step 2.3. To accommodate the increased polygon complexity, the measurement spacing of the two-dimensional ground coordinate points can be reduced from 5 meters to 2 meters to ensure that the lines connecting vertices do not excessively cross the actual boundaries.

[0029] In another optional implementation, while maintaining the global boundaries of 2 seconds and 60 seconds for the expected travel time lower and upper limits, the limits are set separately for different channels based on their measured lengths. For example, for channels between 30 and 50 meters in length, the upper limit for the expected travel time is set to 40 seconds; for channels between 10 and 30 meters in length, the upper limit is set to 25 seconds; and for channels less than 10 meters in length without turnstiles, the upper limit is set to 10 seconds. The intuitive benefit of this setting is that it combines "path allowable" with "path reasonable": shorter channels are less likely to experience long travel times, thus further compressing the candidate range when establishing cross-camera station candidate edges in step 2.1, reducing subsequent computational load, and improving the reliability of cross-camera station associations.

[0030] In another alternative implementation, in order to improve the stability of the helmet status category under low light and strong backlight conditions, an adaptive enhancement based on local contrast is added after the candidate head area is determined. The intensity of the enhancement is limited to within 20% of the original brightness. By slightly increasing the brightness difference between the foreground and the background, the peak position of color salience is more easily identified, thereby reducing the situation where "wearing in compliance" is misjudged as "bare head".

[0031] In another optional implementation, the generation of the pedestrian bounding box set employs a block-based processing approach in dense crowd scenarios. Specifically, the 1920×1080 frame is divided into 3 columns × 2 rows of sub-blocks. Each sub-block independently performs candidate region generation and non-maximum suppression within its own scope, while maintaining the overlap area ratio threshold at 0.50. Subsequently, non-maximum suppression is performed across sub-block boundaries at the sub-block boundaries to merge cross-boundary candidates. This processing avoids the phenomenon of "high-scoring candidates covering the entire frame while low-scoring candidates are completely suppressed" in large-scale scenes, resulting in a more uniform coverage of the image by the pedestrian bounding box set.

[0032] In another optional implementation, the recognition period remains 1 second, but to cope with network jitter, each camera position is allowed to miss no more than 3 frames within the recognition period. When frames are missing, the human bounding box generator and the helmet state generator process the available frames and output the pedestrian bounding box set and helmet state category as usual. The ReID encoder generates appearance description vectors for the available frames. At the end of the recognition period, the detection set is constructed based on the actual received frame results, without interpolation. This definition ensures that the system maintains stable output during network fluctuations and avoids artifacts caused by frame supplementation from entering subsequent trajectory maps.

[0033] refer to Figure 3Using a pedestrian bounding box as input, with its width labeled W and total height labeled H, the head candidate region is determined by taking a height ratio range from the top boundary downwards within the pedestrian bounding box. Specifically, the head candidate region is a rectangular area between 0.35H and 0.55H from the top of the pedestrian bounding box to its height. This area is labeled with a dashed rectangle, with the same width as the pedestrian bounding box (W) and a height of 0.20H (i.e., 0.55H minus 0.35H). Figure 3 The vertical range of this area is marked with a double-headed arrow as "0.35H-0.55H". The technical reason for selecting this height ratio range is that the posture changes of pedestrians in the image are mainly torso swings, while the head's vertical position within the bounding box is relatively stable. If the candidate area is set too high (e.g., only 0.0H to 0.30H), the head will fall outside the candidate area when the pedestrian looks down, leading to missed detections; if the candidate area is set too low (e.g., 0.60H to 0.80H), it will include interfering areas such as shoulders and arms, reducing the accuracy of the judgment. The range of 0.35H to 0.55H can simultaneously cover postures of looking down, looking up, and slight tilting, and effectively avoid interference caused by arms and tools. Within the head candidate area, the safety helmet detection is based on two features: edge closure and color salience. Figure 3 As shown on the left, an elliptical region, labeled "Helmet," is drawn at the center of the head candidate area. This ellipse represents the detected outer contour of the helmet. The outer edge of the ellipse is marked with a dashed circle, indicating the measurement range of edge closure. The label "closure ratio ≥ 0.70" indicates that when the closure ratio of the ellipse's edge reaches 0.70 or higher, the shape constraint condition is met. The edge closure is calculated as follows: edge contours are aggregated within the head candidate area, and regions with arc-shaped closures are detected. The proportion of the continuous edge portion of this region to the entire contour perimeter is calculated. When this proportion is not less than 0.70, a helmet-shaped object is determined to exist. The advantage of using closure constraints is that even with reflections or bright backgrounds, the arc-shaped closure feature can still exclude these interferences.

[0034] Step 2: Execute the online inference process of topology-gated three-channel GraphTransformer and ReID collaborative consistency decoding, including: Step 2.1: Construct the trajectory map; When constructing the trajectory map, the detection set within one recognition cycle is used as input. Detections of the same target from the same camera position in adjacent frames are aggregated into groups of 3 frames. A detection segment is formed when the overlap area ratio between any two bounding boxes in the 3 frames reaches 0.70 or higher. Using 3-frame aggregation significantly reduces interference from occasional jitter and short-term occlusion while maintaining temporal accuracy: with an input of 25 frames per second, the time span corresponding to 3 frames is approximately 120 milliseconds. The posture changes of a person walking through the factory area are basically continuous within this span. An overlap area ratio of 0.70 ensures that the center of the bounding box remains within the stable part of the same person, thus obtaining a spatiotemporally coherent detection segment. A trajectory map is built using the detection segments as nodes, and three types of directed edges are established. Category A consists of adjacent edges from the same camera position, established by the criterion of a maximum time interval of 1 second for detected segments and a maximum displacement of 100 pixels for the node center point. This is because, at a 1920×1080 resolution, the displacement of a human body within the same camera position's frame at a normal walking speed is typically less than 100 pixels per second; exceeding this value often indicates cross-region switching or false detection. Category B consists of candidate edges from different camera positions, established by the criterion that only when a corresponding directed edge exists from the starting camera position to the ending camera position in the cross-camera topology map, and the time difference between two detected segments lies between the lower limit of the expected traversal time of this directed edge (2 seconds) and the upper limit (60 seconds). This constraint combines the verification of "visible" and "walkable" edges, avoiding false connections to unconnected fields of view. Type C consists of edges that skip edges within the same camera position. The criteria for establishing these edges are a maximum time interval of 3 seconds and a maximum displacement of 150 pixels for the node center point. These edges compensate for gaps in the footage caused by short-term screen absence or obstruction by large devices. The combination of 3 seconds and 150 pixels can cover common pauses such as access control verification and turning to give way, while suppressing erroneous stitching caused by excessively long or distant displacements. To control the size of the trajectory map, within the same recognition cycle, only one directed edge from Type A and one from Type C that earliest meet the criteria are retained for each node, and the first three directed edges from Type B are truncated based on their time difference from smallest to largest. These restrictions ensure that the number of nodes and directed edges in the trajectory map grows linearly during busy shifts, facilitating stable processing within a single recognition cycle.

[0035] refer to Figure 1 , Figure 1 This is a schematic diagram of the cross-camera topology, illustrating the topological connections and path constraints between multiple camera positions within a power plant area. For example... Figure 1As shown, the cross-camera topology diagram contains 9 nodes and 10 directed edges. Each node corresponds to a camera position, named using the prefix "CAM-" followed by a 4-digit incrementing number. Specifically, the diagram shows 9 camera position nodes: CAM-0001, CAM-0002, CAM-0003, CAM-0004, CAM-0005, CAM-0006, CAM-0007, CAM-0008, and CAM-0009. These nodes are arranged in three spatial layers: the first layer contains four camera positions (CAM-0001 to CAM-0004) arranged sequentially along the east-west direction, corresponding to the main passageway of the park; the second layer contains three camera positions (CAM-0005, CAM-0006, and CAM-0007) located below the first layer, corresponding to secondary passage areas; the third layer contains two camera positions (CAM-0008 and CAM-0009) located at the bottom, corresponding to the deep work area. Each node contains multiple attribute information. Taking CAM-0001 as an example, its entry direction code is labeled 1 (North), and its exit direction code is labeled 2 (East), indicating that pedestrians enter the field of view of this camera position from the north and exit from the east. The set of direction code values ​​is {1, 2, 3, 4}, corresponding one-to-one with {North, East, South, West}, set according to the convention of north being up and east being right in the park's overall plan. This unified setting of direction codes facilitates cross-regional path splicing, allowing camera positions in different construction areas to be mapped to a consistent global reference direction. Taking CAM-0004 as an example, its entry direction code is labeled 4 (West), and its exit direction code is labeled 3 (South), indicating that pedestrians enter from the west and exit from the south. Directed edges represent the passable path relationships between adjacent camera positions. The figure shows 10 directed edges, divided into solid lines and dashed lines. Solid lines represent normal travel paths. For example, a directed edge from CAM-0001 to CAM-0002, labeled "2-60s," indicates that the estimated travel time is between 2 seconds and 60 seconds. The label "R001" indicates that the travel area code for this path is R001. Similarly, a directed edge from CAM-0002 to CAM-0003, labeled "2-40s" and "R001," indicates that the estimated travel time for this path is adjusted to a maximum of 40 seconds, but it still belongs to the R001 travel area. A directed edge from CAM-0003 to CAM-0004, labeled "2-25s" and "R002," indicates that the estimated travel time for this path is further adjusted to a maximum of 25 seconds, and the travel area is switched to R002. Dashed lines represent cross-level travel paths, typically corresponding to intersecting passages within the park or cross-area connections. For example, the dashed line from CAM-0002 to CAM-0006 is marked with "2-60s" and "R006", indicating that the expected travel time for this cross-level path is between 2 and 60 seconds, and the passage area code is R006.Similarly, the dashed line from CAM-0003 to CAM-0007 is marked "2-60s" and "R007". The lower and upper limits of the estimated travel time are set in increments based on the actual length of the passage. For example... Figure 1 As shown, for passageways between 30 and 50 meters in length, the estimated maximum travel time is set at 40 seconds, such as CAM-0001 to CAM-0002 and CAM-0002 to CAM-0003; for passageways between 10 and 30 meters in length, the estimated maximum travel time is set at 25 seconds, such as CAM-0003 to CAM-0004; for passageways exceeding 50 meters in length or those requiring access control verification, the estimated maximum travel time remains at 60 seconds, such as cross-floor paths. A lower limit is uniformly set at 2 seconds to exclude obviously impossible human movement and suppress cross-camera misconnections caused by image jitter or false detections. The passageway coding sequence is listed in the passageway order from the starting camera position to the ending camera position. For example, if the path travels from CAM-0001 through CAM-0002 to CAM-0003, the complete access area coding sequence is [R001, R001], indicating that the entire path is within the same administrative area. If the path travels from CAM-0002 through CAM-0003 to CAM-0004, the access area coding sequence is [R001, R002], indicating that the path crosses two administrative areas. The access area coding sequence is used to confirm whether the path traverses the same set of administrative areas during subsequent topology reachability comparisons, thus achieving the separation constraint between "visible" and "passable".

[0036] In one alternative implementation, during trajectory map construction, to address image jitter caused by strong winds or vibrations, a bounding box smoothing process is introduced before forming the detection segment: the median time value of the bounding box center position of a single target within 3 frames is taken, and the median time value of the width and height are taken separately, and then the overlap area ratio is calculated to see if it reaches 0.70. This approach can reduce the impact of single-frame spike noise on the overlap area ratio without changing the 3-frame aggregation and the 0.70 threshold, making the detection segment more stable. Regarding the number of directed edges extracted from Class A and Class C within the same recognition period, the number of edges retained for each class can be increased from 1 to 2 during busy shifts to address multipath phenomena caused by parallel traffic and neighbor occlusion. To avoid the graph size growing too rapidly, the total number of outgoing edges for each node can be limited to 5.

[0037] Step 2.2: Perform three-channel feature injection.

[0038] During the three-channel feature injection, three sets of information—appearance description, temporal description, and topological description—are injected into each node and directed edge of the trajectory map. The appearance description comes from the 256-length appearance description vector obtained in step 1, and the values ​​of each dimension are scaled to the range of 0 to 1, so that the values ​​under different camera positions and lighting conditions are on the same dimension. The direct benefit of this processing is that the clothing texture and color blocks of the same person are still comparable in cross-camera situations, and the numerical scale will not drift due to differences in device gain. The temporal description includes the start timestamp, end timestamp, and duration of the detection segment, and the duration is mapped to four discrete intervals: {0 to 1 second, 1 to 3 seconds, 3 to 10 seconds, and 10 seconds and above}, which are encoded into four mutually exclusive indicators. This segmentation matches the rhythm of actual actions in the factory area: 0 to 1 seconds corresponds to continuous gait while walking, 1 to 3 seconds covers short pauses and sidestepping, 3 to 10 seconds covers access control verification and handover confirmation, and 10 seconds and above corresponds to abnormal dwelling, providing clear temporal granularity for subsequent temporal channels. The topology description includes the gate number, the directed edge identifier across gates, the discrete codes corresponding to the lower and upper limits of the estimated travel time, and the passage area coding sequence. When injected into a type B directed edge, the identifier of that directed edge and the passage area coding sequence are directly appended. When injected into type A and type C directed edges, only the gate number and the first item of the passage area coding sequence are appended, ensuring that the spatiotemporal continuity of the same gate remains distinguishable in the topology channel. To ensure consistency in subsequent comparisons, the passage area coding sequence maintains the order from the starting gate to the ending gate, without lexicographical rearrangement, to avoid different paths being incorrectly treated as the same path in the coding.

[0039] In one optional implementation, when performing three-channel feature injection, the discrete interval of duration remains unchanged, but for the type B directed edges across camera positions, the entry direction code of the starting camera position and the departure direction code of the ending camera position are additionally injected, thereby filtering candidate edges traveling in the opposite direction earlier in the subsequent processing of the topology channel. At the intersection of the trestle and the channel in the factory area, the camera is installed at a higher angle, and the introduction of the direction code can quickly eliminate candidates with "overlapping fields of view but opposite directions", reducing the number of low-priority edges and concentrating effective candidates on the actual passage path. Regarding the scaling method of appearance description, at the boundary between outdoor and indoor areas with extremely large differences in lighting, the appearance description can first be scaled within the camera position, and then a global scaling can be performed at the recognition cycle level to avoid the loss of contrast caused by a single scale during extreme lighting changes.

[0040] Step 2.3: Perform topology-gated three-channel GraphTransformer encoding.

[0041] When performing topology-gated three-channel GraphTransformer encoding, the appearance channel, time channel, and topology channel are executed sequentially on the trajectory graph, with each channel containing two encoder layers. The appearance channel aggregates appearance descriptions on directed edges of types A and C, updating the context of appearance descriptions for short, continuous segments within the same camera position. This local-to-global order allows single-frame appearance information affected by strong light, smoke, or reflections to be corrected by stable information from adjacent frames, thus forming a more robust appearance context within the same camera position. The time channel propagates information along directed edges of types A, B, and C using chronological order as the key, prioritizing earlier preceding segments and giving stronger influence to more recent preceding segments. This strategy directly corresponds to the continuity of human movement; segments closer to the current time are more likely to belong to the same person, thus highlighting the "coherence" attribute of the time context. The topology channel introduces a topology gating system, determining for each directed edge of type B whether it has a cross-camera topology connection. Figure 1 The topology gating system determines the priority of a type B directed edge. When all three conditions are met—directed connectivity, the time difference between two detected segments falling within the expected travel time interval, and the consistency of the encoded sequences of the travel areas—the edge is marked as a high-priority edge. Edges that do not meet all three conditions are marked as low-priority edges. During encoding, the topology gating system retains messages from high-priority edges and includes them in node updates, while skipping messages from low-priority edges at this layer. This gating eliminates the influence of unreasonable paths during the encoding phase, ensuring that the joint context representation is derived primarily from edges that are "visible, traversable, and time-reachable." After executing two layers sequentially across the three channels, a joint context representation is generated for each node. This representation includes corrected appearance information within the same camera position, time information focused sequentially, and cross-camera connectivity information filtered by the topology gating system. This information is used for subsequent set-growth identity grouping and event matching.

[0042] refer to Figure 2The process executes sequentially from left to right. The input is a type B cross-camera edge, which includes the starting camera number and the ending camera number. In the example in the figure, the starting point of the input edge is CAM-0001, and the ending point is CAM-0002. The first decision box is "Constraint 1: Topological Connectivity," which is marked "Does a directed edge exist in the cross-camera topology?". This indicates whether there is a corresponding directed edge from the starting camera CAM-0001 to the ending camera CAM-0002 in the cross-camera topology. This constraint is used to verify whether the two cameras are physically connected, i.e., whether "can see" corresponds to "can walk." If there is no corresponding directed edge in the cross-camera topology, it means that although the two cameras may have overlapping fields of view, there is actually no passable path, and this type B edge should not be adopted. When the decision result is "yes," the process proceeds to the second decision box, "Constraint 2: Time Interval." The box labeled "Lower limit ≤ Δt ≤ Upper limit?" and "2s ≤ Δt ≤ 60s" indicates whether the time difference Δt between the two detection segments of the B-type edge is within the expected traversal time interval. Specifically, the lower and upper limits of the expected traversal time for the directed edge are queried from the cross-camera topology graph (examples in the figure are 2 seconds and 60 seconds), and then it is determined whether the actual measured time difference Δt satisfies 2 seconds ≤ Δt ≤ 60 seconds. This constraint is used to verify whether the cross-camera connection is reasonable in time, excluding edges with time differences that are too short (suspected false detection) or too long (suspected breakpoint). When the determination result is "yes", the process proceeds to the third determination box "Constraint 3: Area Coding". The box labeled "Travel area coding sequence consistent?" indicates whether the travel area coding sequence of the B-type edge is completely consistent with the travel area coding sequence of the corresponding directed edge in the cross-camera topology graph. The travel area coding sequence maintains the order from the starting camera position to the ending camera position, without lexicographical rearrangement. This constraint is used to verify whether a path traverses the same group of management areas, avoiding false connections where paths are "visible but cannot be traversed." Below the third decision box is the final decision box, "Three-Constraint Decision," marked with "All Satisfied?", indicating a summary of the decision results for the aforementioned three constraints. When all three constraints—topological connectivity, time interval, and area coding—are satisfied, the decision result is "Yes," and the process moves to the right into the "High-Priority Edge" processing branch; when any one or more of the three constraints are not satisfied, the decision result is "No," and the process moves down into the "Low-Priority Edge" processing branch.

[0043] In an optional implementation, when performing topology-gated three-channel GraphTransformer encoding, each channel remains at two layers, but a small-scale path expansion check is added to the second layer of the topology channel: when a type B directed edge is marked as a low-priority edge because its time difference slightly exceeds the upper limit of the expected traversal time, if there are intermediate positions of the same passage area encoding sequence near both ends of the edge, and inserting such intermediate positions can compress the time difference of the two short edges into the expected traversal time interval, then an alternative path consisting of two high-priority edges is temporarily constructed during the encoding stage to participate in node updates; this processing provides a compensation channel for actual passage in critical situations without changing the predetermined upper limit, which is common in scenarios where access control queuing causes the time difference to be lengthened but the actual walking path is still reasonable.

[0044] Step 2.4: Construct an identity consistency decoder.

[0045] When constructing the identity consistency decoder, the joint context representation obtained in step 2.3 is used as input to perform set-growing identity grouping. First, all nodes within a recognition period are sorted from high to low according to their appearance similarity scores, and the node with the highest score is selected as the anchor point to create an identity group. The reason for choosing appearance similarity score as the primary sorting criterion is that appearance description is the most effective way to distinguish between clothing and equipment differences within the same recognition period, and it is independent of time and topological constraints. Determining the anchor point based on appearance first can reduce subsequent backtracking. After creating the identity group, candidate nodes are selected from the successor nodes of the directed edges of type A, type B, and type C connected to the anchor point, and three criteria are applied sequentially for inclusion: appearance similarity score of 0.82 or higher; increasing time sequence with time intervals within the same camera position not exceeding 1 second and cross-camera time intervals within the expected traversal time interval of the corresponding directed edge; and a directed path from the previous node position to the candidate node position exists on the cross-camera topology map with a path hop count not exceeding 8. The approach prioritizes appearance, then time, and finally topology. This prioritizes maintaining appearance and temporal continuity while ensuring traversability. The combined constraint of these two factors significantly reduces false inclusions caused by appearance fluctuations due to momentary occlusion or backlighting. When multiple candidate nodes simultaneously satisfy all three criteria, the candidate node with the earlier timestamp is selected, followed by the candidate node with the higher appearance similarity score. This decision balances the empirical rule that "whoever appears first is more likely to be the same person" with the judgment rule that "those with closer appearances are more credible." After an identity group absorbs a new node, that node is used as the new previous node, and candidate nodes are continuously selected from its outgoing edges until no candidate node satisfies the above three criteria or the recognition period ends. This yields a time-sorted camera position sequence and a time sequence, while simultaneously generating cross-camera IDs. After completing an identity group, the node with the highest appearance similarity score is selected again from the remaining unassigned nodes as the anchor point, and the above process is repeated until all nodes within the recognition period are assigned or cannot meet the absorption criteria. This process does not rely on historical data; all judgments are based on the joint context representation and cross-camera topology map within the current recognition period, making online processing deterministic and verifiable.

[0046] In an optional implementation, when constructing the identity consistency decoder, the anchor point selection is changed from "the highest appearance similarity score of a single node" to "the highest temporal median of appearance similarity scores within a 3-frame time window for the same camera position." When instantaneous reflections or smoke cause an anomaly in a frame score, the temporal median can suppress the impact of outliers, making the anchor point more stable. This method still maintains the 0.82 threshold and the path hop count not exceeding 8.

[0047] Step 2.5: Perform event decoder matching with template.

[0048] When performing event decoder and template matching, for each completed identity group, event matching is performed based on the camera position sequence, time sequence, and helmet status category sequence. Event templates include Template A, Template B, and Template C. Template A is for bare-helmet penetration, determined by the identity group's helmet status category being bare-helmet at the first camera position, followed by at least four camera position changes, forming a sequence of five consecutive camera positions, with a cumulative dwell time of 20 seconds or more. The requirement of being bare-helmet at the first camera position clearly identifies the starting point of the violation, avoiding misjudging short-term helmet removal as penetration. The at least four camera position changes and the 20-second threshold together limit the scope of "penetration," distinguishing between short-distance tentative entry and genuine long-distance passage. Template B is for excessive stay in high-voltage zones, determined by the identity group's continuous stay time of 20 seconds or more within a set of camera positions marked with high-voltage zone markers. High-voltage zone markers are indicated by the camera position attribute field "isHV" being 1. The reason for setting continuous stays instead of segmented accumulation is that the risk in high-voltage areas comes from continuous exposure, and a single continuous stay exceeding 20 seconds is more in line with safety management requirements. Template C is reverse passage. The judgment condition is that the identity group completes one camera position switch along a directed edge in the opposite direction in the cross-camera topology map, and the two ends of the switch are consistent with the start and end points of the corresponding directed edge, and the entry and exit direction codes are consistent with the reverse movement. The advantage of using direction codes for simultaneous verification is that it can distinguish between scenes with overlapping fields of view but consistent actual passage directions and truly reverse passage scenes, reducing false alarms. The event decoder generates an event record for identity groups that meet any template. The event record includes the event type code, the associated cross-camera ID, the camera position sequence, and the start and end timestamps. To make the records traceable, each event record includes the camera position index and the corresponding safety helmet status category or direction code value used to trigger the judgment, which facilitates quick playback and verification during review.

[0049] In one optional implementation, when the event decoder matches the template, the cumulative dwell time threshold for template A is set to 25 seconds near the open-air coal yard and the plant's outer boundary, and 20 seconds in the main passageways within the plant. Since walking paths are longer in open-air areas, the time people spend on the road naturally increases. Raising the threshold to 25 seconds reduces the probability of normal passage being misjudged as continuous patrol, while maintaining 20 seconds in corridor scenarios preserves sensitivity to rapid movement. All other judgment conditions remain unchanged.

[0050] Step 2.6: Generate collaborative and consistent joint optimization decisions.

[0051] When generating collaborative consistency joint optimization decisions, two rounds of collaborative consistency calibration are performed within the same forward link. Round 1 primarily uses the cross-camera ID and camera position sequence produced by the identity consistency decoder to first generate initial event records, and then triggers the path completer. The path completer checks for gaps between the camera position sequence and time sequence of the identity group. When the time interval between two adjacent camera position switches does not exceed 15 seconds and there is an intermediate camera position with a consistent access area encoding sequence between the two cameras in the cross-camera topology map, it attempts to insert this intermediate camera position as a transition point, forming an alternative path composed of two type B directed edges. Before insertion, the expected traversal time interval and access area encoding sequence are checked for the two alternative edges respectively. Insertion is only completed when both conditions are met. This approach can repair segment interruptions caused by temporary occlusion, access control queuing, or crowd occlusion without relaxing the global threshold, completing the real access path into a reachable path. After insertion, steps 2.3 and 2.4 are returned to perform one incremental update to ensure that the joint context representation remains consistent with the identity group within the current recognition period. Round 2 primarily uses the event records generated by the event decoder to calibrate the camera position sequence for the identity groups that triggered Template A and Template C. For Template A, it checks whether the camera position switching corresponds one-to-one with the directed edges of Type B. If a switching is found to lack a corresponding directed edge of Type B but there is a skip connection caused by a short jump within the same camera position, a directed edge of Type C with a time interval not exceeding 3 seconds and a node center point displacement not exceeding 150 pixels is inserted at that position, and the skip connection is split into two segments to ensure that the camera position sequence is consistent with the cross-camera topology map. For Template C, it verifies whether the entry and exit direction codes of the switching for reverse passage are consistent with the reverse walking. If they are inconsistent, the triggering of that template is canceled and marked as pending verification to avoid false alarms caused by incorrect camera installation orientation.

[0052] To clarify the convergence criteria, edge scores are assigned to each directed edge within an identity group and statistically analyzed over the recognition period. The edge scores employ a three-channel consistency hit scoring method: an edge score is set to 1.00 when all three constraints—topological connectivity, time interval, and access area coding sequence—are met; 0.67 when any two are met; 0.33 when any one is met; and 0.00 when none are met. The edge scores of all identity groups within the recognition period are averaged to obtain the average edge score. The inference process for that recognition period ends when the average edge score reaches 0.80 or higher, or when two rounds of calibration have been completed, and a cross-camera ID and a global list of insecure behavior events are generated simultaneously. This segmented scoring method, rather than continuous weighting, ensures that the convergence conditions and judgment criteria remain interpretable and verifiable. The source of any edge score can be mapped to a specific hit of the three constraints, facilitating verification of each edge during security audits.

[0053] Step 3: Online publishing and linkage, including outputting structured results after each recognition cycle and sending them to the scheduling end via message bus to trigger linkage strategies.

[0054] In one implementation, after a recognition cycle ends, the system generates a structured result. The structured result is organized as an ordered set of key-value pairs. Fields include the recognition cycle number, start timestamp, end timestamp, cross-camera ID string, camera position sequence, entry and exit times for each camera position, a global list of unsafe behavior events and event type codes, a summary of the average edge score, and a segment index for playback. The recognition cycle number increments from 0 for each day to align with the daily log at the scheduling end; the start and end timestamps are accurate to milliseconds to ensure seamless concatenation with adjacent recognition cycles; the cross-camera ID string is fixed at 16 characters and uses a set containing only uppercase letters and numbers for easy verbal verification; the camera position sequence is represented as a chronologically ordered list of camera position numbers; the entry and exit times for each camera position are accurate to milliseconds for easy camera-by-camera playback on a large screen; and the global list of unsafe behavior events is a set of records arranged in order of triggering. Each record includes an event type code, the triggering camera position index location, the associated safety helmet status category or direction code value, and the event start and end timestamps; the average edge score summary is the average edge score of all identity groups within the identification period, with precision retained to two decimal places; the segment index for playback includes a playback start and end time suggestion for each camera position switch, with the suggested start point taken as 5 seconds before the switch and the suggested end point taken as 5 seconds after the switch. This time boundary can cover both the approach process before entry and the confirmation process after departure, facilitating manual review.

[0055] After the structured results are generated, the release preparation phase begins, involving sorting, merging, and deduplication. Sorting prioritizes event severity, in the order of template A, template C, and template B. This arrangement is based on the fact that unmarked camera penetration has the most direct impact on park production safety, followed by reverse passage, and lastly, excessive timeouts in high-voltage areas. Merging employs a time window strategy: multiple records from the same cross-camera ID with the same event type code within a 3-second window are merged into one. The earliest timestamp within the window is used as the event start point, and the latest timestamp as the event end point. The merged record carries the set of camera position indexes involved within the window. This merging strategy can suppress alarm storms when the same person passes through multiple camera positions consecutively. Deduplication uses a unique key strategy: the unique key is the identification cycle number plus the cross-camera ID plus the event type code plus the camera position index. For records with duplicate unique keys, only the first one generated is retained. This key design ensures that each record corresponds one-to-one with a specific replayable switching location, facilitating auditing.

[0056] It is sent through the message bus at the time of release. Each piece of sent content consists of a message header and a message body. The message header contains the message unique number, cross-camera ID, event type code, priority, expiration timestamp, sequence key, and integrity check code. The length of the message unique number is 20 characters, which is generated by concatenating the timestamp and a random number; the priority value is at level 3, corresponding to the aforementioned sorting rules; the expiration timestamp is set to 3 seconds after the message is generated. Messages older than 3 seconds have reduced significance for linkage and are discarded by the scheduling end after expiration; the sequence key takes the cross-camera ID to ensure that multiple messages of the same person are processed in the same order on the message bus and the scheduling end; the integrity check code is 16 bytes long, covering the original message body, facilitating quick verification of whether tampering has occurred in the link. The message body is a trimmed subset of the structured result, retaining the cross-camera ID string, camera position sequence, global list of unsafe behavior events and event type codes, and fragment index for playback. The sending adopts an at-least-once delivery strategy, with a retry interval of 200 milliseconds and a maximum retry count of 2 times; the combination of 200 milliseconds and 2 times is chosen considering that transient congestion in the园区 network usually eases within hundreds of milliseconds, and excessive retries would occupy the message bus bandwidth. Wait for the scheduling end to return a receipt after each send. The receipt contains the message unique number and the reception timestamp. When the reception timestamp is earlier than the message expiration timestamp, it is marked as successful. The queue depth of the message bus is set to 1000 messages or 5 megabytes. When either of the two reaches the threshold, congestion control is entered. During congestion control, only messages with the highest priority are queued, and the remaining messages wait for the next send within the current recognition cycle; adopting "priority first" congestion control can ensure that critical alarms reach the scheduling end first during extremely busy periods.

[0057] It should be noted that the text contains some placeholders like "园区 network" which may need to be further specified according to the actual context. If this is a specific industry term with a more accurate translation, it should be adjusted accordingly.Upon receiving a message, the dispatcher immediately triggers the linkage strategy. The strategy configuration is maintained in a mapping table, with the key being the event type code and the value being the linkage action sequence. For records corresponding to template A of the event type code, the dispatcher pushes an alarm to the gatekeeper and adjacent security terminals. The alarm includes a cross-camera ID string, camera position sequence, event start and end timestamps, and a segment index for playback. Playback is triggered on the large screen in camera position sequence order. The playback start point is the suggested start point of the first camera position, and the playback end point is the suggested end point of the last camera position. During playback, the cross-camera ID string and the current camera position number are displayed in the lower right corner of the screen. The frame rate is set to 15 frames per second. A frame rate lower than the original frame rate reduces the overhead on the link during playback while preserving the continuity of actions, facilitating quick confirmation by on-duty personnel. Pop-up windows on the gatekeeper and adjacent security terminals provide "Intercept and Record Arrival Time" and "Continue Observation" buttons. Both buttons write the operation timestamp to the audit log, forming a traceable closed loop. For records corresponding to template B in the event type code, the dispatch terminal highlights the set of stations marked with high-voltage zone symbols on the park's overhead view on the large screen, and overlays a countdown indicator in this area. When the continuous stay reaches 40 seconds, the indicator color is increased to a more prominent color and an update is pushed to the team leader's terminal. The two-level duration indicator can distinguish between "short stays due to normal verification" and "long stays due to increased risk". For records corresponding to template C in the event type code, the dispatch terminal triggers an audio-visual prompt in the status bar of the corresponding station and overlays an arrow direction indicator on the large screen. The direction is determined by the entry and exit direction codes, making it easy for inspection personnel to identify the direction of the reverse path at a glance.

[0058] To meet the end-to-end processing latency target of 1 second or less, this implementation adopts a fixed budget allocation: structured results generation, sorting, merging, and deduplication are completed within 100 milliseconds after the identification cycle ends; message bus sending and acknowledgment are then completed within 150 milliseconds; and the scheduling end completes message parsing, linkage strategy matching, alarm push, and playback initiation within 600 milliseconds. This allocation divides 1 second into 3 segments, with the most time-consuming segment allocated to the visualization and push process on the scheduling end. This ensures small and stable time slices for both the identification and link sides, making it easier to meet the overall target. Each step records its processing time locally, and when the budget is exceeded, a "duration alarm flag" is added to the structured results. The scheduling end displays a small icon in the corner of the large screen to remind operators to monitor the link health.

[0059] To ensure audit traceability and cross-system consistency, the dispatch terminal writes the original message, linkage action, and on-duty personnel operation records to the audit storage after completing the linkage. Each record in the audit storage includes a unique message ID, cross-camera ID string, event type code, camera position sequence, event start and end timestamps, average edge score summary, receipt receipt timestamp, push completion timestamp, and a set of on-duty personnel operation timestamps. A retention period of 90 days is recommended. On-duty personnel can retrieve records by unique message ID or cross-camera ID string and replay the corresponding segment index at any time to verify the judgment process. The audit storage and message bus use different write paths to avoid mutual interference during peak periods.

[0060] In one optional implementation, the structured result is appended with a sequence of thumbnails for rapid confirmation. The thumbnail sequence samples one frame from each camera position in the camera position sequence, with the sampling time set 200 milliseconds after the camera's entry time. The cropping range is calculated by extending the margins of the pedestrian bounding box by 10% on each side. It is then scaled to 320 x 180 pixels and compressed to the Joint Image Experts Group (JEP) format with a compression quality factor of 80. Choosing a sampling time of 200 milliseconds after entry avoids edge blurring immediately upon entering the frame; extending the margins by 10% preserves the edges of the safety helmet and clothing texture; and scaling to 320 x 180 ensures immediate display in the pop-up windows of the gatekeeper and adjacent security terminals while maintaining manageable network transmission load. The thumbnail sequence is an optional field in the message body; when the message bus is under congestion control, the text field is carried first, and the thumbnail sequence is only appended when the link is idle.

[0061] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative examples, and those skilled in the art can omit, substitute, and change the details of the above methods and systems in various ways without departing from the principles and essence of the present invention.

Claims

1. A method for online identification of unsafe human behavior based on spatiotemporal Transformer, characterized in that, The method includes: Step 1: Scene setting and cross-camera topology map solidification, including establishing node records and directed edge records, and connecting video streams to generate pedestrian bounding box sets, safety helmet status categories and appearance description vectors, forming a detection set for a recognition cycle; Step 2: Execute the online inference process of topology-gated three-channel GraphTransformer and ReID collaborative consistency decoding, including: Step 2.1: Construct the trajectory map; Step 2.2: Perform three-channel feature injection; Step 2.3: Perform topology-gated three-channel GraphTransformer encoding; Step 2.4: Construct an identity consistency decoder; Step 2.5: Perform event decoder matching with template; Step 2.6: Generate collaborative and consistent joint optimization decisions; Step 3: Online publishing and linkage, including outputting structured results after each recognition cycle and sending them to the scheduling end via message bus to trigger linkage strategies.

2. The method as defined in claim 1, characterized in that, Step 1, establishing node records and directed edge records, includes: creating a node record for each camera position. This node record contains the camera position number, the two-dimensional ground coordinate point, at least 8 vertices of the field of view polygon outline, and one entry direction code and one exit direction code. The value sets for the entry and exit direction codes are as follows: One-to-one correspondence; a directed edge record is established based on the connectivity between the park's fixed access channels, access gates, and fence gates. This directed edge record includes the starting point gate number, the ending point gate number, the estimated travel time (lower limit 2 seconds and upper limit 60 seconds), and the access area code sequence.

3. The method as defined in claim 1, characterized in that, Step 1 involves accessing the video stream to generate a set of pedestrian bounding boxes, helmet state categories, and appearance description vectors, forming a detection set for a recognition cycle. This includes: accessing a 1920x1080 resolution video stream at 25 frames per second; executing a human bounding box generator on each frame to output a set of pedestrian bounding boxes, and performing non-maximum suppression based on an overlap area ratio of 0.50 or higher; and executing a helmet state generator to output helmet state categories, where the set of values ​​for the helmet state categories is... ; Execute a ReID encoder to output an appearance description vector with a length of 256; Collect the human body frame, safety helmet status category and appearance description vector of each camera position in a recognition cycle of 1 second, and form the detection set of this recognition cycle.

4. The method as defined in claim 1, characterized in that, Step 2.1: Trajectory map construction includes: within each recognition cycle, the detection of the same target from the same camera position in adjacent frames is aggregated into groups of 3 frames. When the overlap area ratio between any two bounding boxes of the 3 frames reaches 0.70 or above, a detection segment is formed. A trajectory map is built with the detection segment as the node, and 3 types of directed edges are established: Type A: Time-adjacent edges from the same camera position, established with a time interval limit of 1 second and a node center point displacement limit of 100 pixels; Type B: Cross-camera candidate edges, established with the criterion that there is only a corresponding directed edge from the starting camera position to the ending camera position in the cross-camera topology map and the time difference between the two detection segments is within the expected travel time interval of the corresponding directed edge; Type C: Jump edges from the same camera position, established with a time interval limit of 3 seconds and a node center point displacement limit of 150 pixels.

5. The method as defined in claim 4, characterized in that, Step 2.2: Perform three-channel feature injection, including injecting three sets of descriptions for each node and edge of the trajectory graph: appearance description, which is the normalized sequence of the 256-dimensional appearance description vector from the ReID encoder; and time description, which is the start time stamp, end time stamp, and duration of the detected segment, with the duration divided into intervals. Perform discrete coding; topological description, obtain discrete codes for gate number, cross-gate directed edge identifier, lower and upper limits of expected travel time, and travel area coding sequence.

6. The method as defined in claim 5, characterized in that, Step 2.3: Perform topology-gated three-channel GraphTransformer encoding, which includes: sequentially executing attention encoder stacks for the appearance, time, and topology channels on the trajectory graph. Each channel contains two encoder layers, and each layer sequentially performs multi-head attention, feedforward update, and layer normalization. The appearance channel aggregates appearance descriptions on directed edges of types A and C to form a local appearance context. The time channel performs message propagation on directed edges of types A, B, and C with the time sequence as the key to form a time context. The topology channel introduces a topology gate, which performs three criteria for each directed edge of type B: there is directed connectivity consistent with the cross-camera topology graph, the time difference between two detection segments is within the expected traversal time interval of the corresponding directed edge, and the encoding sequence of the traversal area is consistent. When all three criteria are met, the directed edge of type B is marked as a high-priority edge, and the directed edge of type B that is not marked as a high-priority edge is marked as a low-priority edge. After the three channels are executed sequentially, a joint context representation is generated for each node.

7. The method as defined in claim 6, characterized in that, Step 2.4: Constructing an identity consistency decoder includes: performing set-growth identity grouping with the joint context representation as input. The process is as follows: create an identity group with the node with the highest score as the anchor point, and absorb new nodes and form cross-camera IDs according to the following three criteria: Criterion 1: the appearance similarity score reaches 0.82 or above; Criterion 2: the time order is increasing and the time interval between the same camera position does not exceed 1 second, and the time interval between cross-camera positions is located in the expected traversal time interval of the corresponding directed edge; Criterion 3: there is a directed path from the previous node position to the candidate node position on the cross-camera topology graph and the number of path hops does not exceed 8; when multiple candidate nodes meet all three criteria at the same time, the candidate node with the earlier timestamp is selected, followed by the candidate node with the higher appearance similarity score; after the identity group is absorbed within this recognition period, a time-sorted camera position sequence and time sequence are formed.

8. The method as defined in claim 7, characterized in that, Step 2.5: Executing the event decoder and template matching includes: For each completed identity group, event matching is performed within the identification period based on the camera position sequence, time sequence, and helmet status category sequence. The event template contains three subcategories with numerical thresholds for determining the criteria: Template A is bare-headed through patrol, the criteria being that the identity group's helmet status category is bare-headed at the first camera position and subsequently completes at least four camera position switches, forming a sequence of continuously passing through five camera positions with a cumulative dwell time of 20 seconds or more; Template B is high-voltage zone timeout stay, the criteria being that the identity group's continuous dwell time within the set of camera positions marked with high-voltage zones reaches 20 seconds or more, and the high-voltage zone marker is represented by the camera position attribute field isHV being 1; Template C is reverse passage, the criteria being that the identity group completes one camera position switch in the opposite direction of a directed edge in the cross-camera topology map, and the two camera positions are consistent with the start and end points of the corresponding directed edge; The event decoder generates an event record for identity groups that satisfy any template. The event record includes the event type code, the associated cross-camera ID, the camera position sequence, and the start and end timestamps.

9. The method as defined in claim 8, characterized in that, Step 2.6: Generating collaborative consistency joint optimization decisions includes: performing two rounds of collaborative consistency calibration within the same forward link; in round 1, the cross-camera ID and camera position sequence generated by the identity consistency decoder are used as the main output for initial event records and triggering a path completer. The path completer inserts virtual type B directed edges for identity groups with camera position sequence gaps and cross-camera topology connectivity within a time interval of no more than 15 seconds and returns to execute steps 2.3 and 2.4 for one incremental update; in round 2, the event records generated by the event decoder are used as the main output for camera position sequence calibration of identity groups triggering templates A and C, so that camera position switching corresponds one-to-one with the directed edges in the cross-camera topology; when the average edge score of all identity groups in this identification period reaches 0.80 or above or two rounds of calibration have been completed, the inference of this identification period ends and a cross-camera ID and a global list of insecure behavior events are generated simultaneously.

10. The method as defined in claim 1, characterized in that, Step 3, online publishing and linkage, includes: outputting a structured result after each recognition cycle, which includes cross-camera ID strings, camera position sequence, entry and exit times of each camera position, a global list of unsafe behavior events, and event type codes; sending it to the dispatch terminal via the message bus; the dispatch terminal triggers the linkage strategy based on the event type code, pushes alarms to the gatekeeper and adjacent security terminals for the record corresponding to template A of the event type code, and calls the playback interface to play back the data in the camera position sequence order; the end-to-end processing delay target is 1 second or less.