Tunnel resident person detection method and system
By collecting and fusing visible light and thermal imaging images of tunnel sections, and using a cross-modal feature reconstruction network to generate enhanced feature maps, accurate detection of personnel stranded in tunnels under smoke and dust environments is achieved, solving the problem of inaccurate detection in existing technologies and realizing automated real-time detection and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中交资产管理有限公司
- Filing Date
- 2026-04-14
- Publication Date
- 2026-06-16
AI Technical Summary
Existing tunnel monitoring systems struggle to accurately detect personnel lingering in dusty environments. Current technologies lack in-depth exploration and effective integration of the complementary characteristics of visible light and thermal imaging information, resulting in an inability to reliably and automatically perform real-time detection and early warning.
Visible light image frame sequences and thermal imaging image frame sequences of the tunnel section are collected. Reflection attenuation feature maps and radiation penetration feature maps are generated through complementary feature excitation operations. Feature compensation and fusion are performed using a cross-modal feature reconstruction network to generate enhanced personnel thermal radiation feature maps and reflection feature maps. Heat source tracking and spatiotemporal dwell patterns are analyzed to generate the marking results of personnel staying in the tunnel.
It effectively overcomes the interference of smoke and dust, achieves accurate detection of people stranded in tunnels, and automatically generates marking results of stranded people, including the start time and duration of their stay.
Smart Images

Figure CN122223657A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tunnel safety technology, and more specifically, to a method and system for detecting people stranded in tunnels. Background Technology
[0002] In semi-enclosed spaces such as tunnels, personnel lingering is a significant risk factor for safety accidents. Existing tunnel monitoring systems largely rely on visible light cameras for video analysis. However, tunnels are often filled with smoke and dust from vehicle exhaust and construction, which severely attenuates visible light, causing visible light-based personnel detection methods to fail or experience a sharp decline in accuracy. While thermal imaging technology can penetrate some smoke and dust to detect heat radiation, its imaging resolution is low and it is easily affected by environmental heat sources, making it difficult to accurately locate and identify static or slowly moving personnel when used alone. Furthermore, existing technologies lack in-depth exploration and effective fusion mechanisms of the complementary characteristics of visible light and thermal imaging information in smoke and dust environments, and they also fail to effectively analyze the key behavioral pattern of lingering from personnel movement trajectories. This results in an inability to reliably and automatically detect and warn of the risk of personnel lingering in tunnels in real time. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for detecting people stranded in tunnels.
[0004] In a first aspect, embodiments of the present invention provide a method for detecting people stranded in a tunnel, the method comprising:
[0005] The visible light image frame sequence and the thermal imaging image frame sequence of the tunnel section are acquired. Each visible light image frame in the visible light image frame sequence and the thermal imaging image frame sequence with the same acquisition time identifier constitute a cross-modal image frame pair.
[0006] Complementary feature excitation operation is performed on the cross-modal image frame pair to obtain the reflection attenuation feature map of the visible light image frame and the radiation penetration feature map of the thermal imaging image frame. The reflection attenuation feature map records the propagation loss distribution of visible light in the tunnel smoke environment, and the radiation penetration feature map records the transmission intensity distribution of thermal radiation in the tunnel smoke environment.
[0007] The reflection attenuation feature map and the radiation penetration feature map are input into a cross-modal feature reconstruction network to perform feature compensation and fusion operations, generating an enhanced personnel thermal radiation feature map and an enhanced personnel reflection feature map;
[0008] Based on the enhanced thermal radiation feature map and the enhanced reflection feature map of the personnel, a heat source tracking operation is performed on the stranded personnel to obtain a set of heat source movement trajectory sequences. The set of heat source movement trajectory sequences includes the spatial coordinate change path of each heat source in a continuous cross-modal image frame pair.
[0009] A spatiotemporal dwelling pattern parsing operation is performed on the set of heat source movement trajectory sequences to generate a tunnel dwelling personnel marking result, which includes a dwelling start time parameter and a dwelling duration parameter.
[0010] Secondly, embodiments of the present invention provide a tunnel personnel detection system, including at least one service node; the service node includes a storage unit and a computing unit; the storage unit is used to store program code; the computing unit is used to run the program code to execute the tunnel personnel detection method described in the first aspect.
[0011] Compared to existing technologies, the beneficial effects provided by this invention include: The method and system for detecting people stranded in tunnels, as disclosed in this invention, include: firstly, acquiring visible light and thermal imaging image sequences of the tunnel section, and pairing them to form cross-modal image frame pairs; extracting the visible light reflection attenuation feature map and the thermal imaging radiation penetration feature map through complementary feature excitation operations; using a cross-modal feature reconstruction network to perform feature compensation and fusion on the two, generating enhanced personnel thermal radiation feature maps and reflection feature maps; performing heat source tracking based on the enhanced feature maps to obtain a heat source movement trajectory sequence; and finally, analyzing the trajectory through a spatiotemporal dwelling mode to generate a tunnel stranded personnel marking result containing the dwelling start time and duration. This invention effectively overcomes smoke and dust interference and achieves accurate detection of people stranded in tunnels. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating the steps of the method for detecting stranded personnel in a tunnel provided in an embodiment of the present invention.
[0014] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0016] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0017] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the method for detecting people stranded in tunnels provided in this embodiment of the present disclosure. The method for detecting people stranded in tunnels will be described in detail below.
[0018] Step S201: Acquire visible light image frame sequence and thermal imaging image frame sequence of tunnel section. Each visible light image frame in the visible light image frame sequence and the thermal imaging image frame with the same acquisition time sequence identifier in the thermal imaging image frame sequence constitute a cross-modal image frame pair.
[0019] Step S202: Perform complementary feature excitation operation on the cross-modal image frame pair to obtain the reflection attenuation feature map of the visible light image frame and the radiation penetration feature map of the thermal imaging image frame. The reflection attenuation feature map records the propagation loss distribution of visible light in the tunnel smoke environment, and the radiation penetration feature map records the transmission intensity distribution of thermal radiation in the tunnel smoke environment.
[0020] Step S203: Input the reflection attenuation feature map and the radiation penetration feature map into the cross-modal feature reconstruction network to perform feature compensation fusion operation, and generate an enhanced personnel thermal radiation feature map and an enhanced personnel reflection feature map;
[0021] Step S204: Perform heat source tracking operation on stranded personnel based on the enhanced personnel thermal radiation feature map and the enhanced personnel reflection feature map to obtain a set of heat source movement trajectory sequences. The set of heat source movement trajectory sequences includes the spatial coordinate change path of each heat source in a continuous cross-modal image frame pair.
[0022] Step S205: Perform a spatiotemporal dwelling mode parsing operation on the heat source movement trajectory sequence set to generate a tunnel dwelling personnel marking result. The tunnel dwelling personnel marking result includes a dwelling start time parameter and a dwelling duration parameter.
[0023] In this embodiment of the invention, for example, firstly, the server controls a dual-mode acquisition device deployed in the tunnel monitoring area to acquire image data. This device integrates a visible light camera and a thermal imaging camera, and has a built-in high-precision synchronization clock. The server sends a synchronization acquisition command to the dual-mode acquisition device, and the device synchronously captures visible light and thermal imaging images of the tunnel section according to the command. For each acquisition action, the server assigns the same, unique acquisition timing identifier (e.g., encoding based on UTC timestamps) to a visible light image frame sequence and a thermal imaging image frame sequence acquired simultaneously. Within the sequence, visible light image frames and thermal imaging image frames with the same acquisition timing identifier are organized internally by the server into a "cross-modal image frame pair." These image frame pairs constitute the basic data unit for all subsequent analysis and processing.
[0024] Subsequently, the server performs a "complementary feature excitation operation" on each cross-modal image frame pair to explore the attenuation characteristics of visible light images and the penetration characteristics of thermal imaging images in smoke and dust environments, respectively. For visible light image frames, the server first extracts their pixel brightness distribution matrix. Next, the server calculates the brightness attenuation gradient value for each pixel in this matrix, reflecting the degree of brightness change relative to its neighborhood; areas with high smoke and dust concentration typically exhibit obvious gradient change patterns. The server arranges the brightness attenuation gradient values of all pixels according to their coordinates, generating a brightness attenuation gradient distribution map. Then, the server performs gradient direction histogram statistics on this distribution map, summing the gradient magnitudes within different gradient direction intervals to generate a feature vector characterizing the overall gradient direction energy distribution of the image. Based on this feature vector, the server can identify specific texture regions in the image caused by smoke and dust scattering, i.e., smoke and dust scattering regions. The server extracts the mean of the brightness attenuation gradient values of all pixels within this region as a parameter to quantify the smoke and dust scattering intensity of the current image. Using this parameter, the server performs a non-linear mapping (i.e., attenuation mapping) on the brightness value of each pixel in the original visible light image, calculating the "reflection attenuation coefficient" for each pixel. This coefficient simulates the theoretical loss of visible light as it reflects from that point and passes through smoke and dust to reach the camera; a lower coefficient indicates more severe attenuation. Finally, the server arranges the reflection attenuation coefficients of all pixels according to their original coordinates to generate a "reflection attenuation feature map."
[0025] For thermal imaging image frames within the same cross-modal image frame pair, the server processes them in parallel. The server extracts the pixel temperature value matrix and calculates the temperature difference between each pixel and its neighbors, forming a temperature gradient field distribution map. The server executes an isotherm tracking algorithm on this distribution map to identify closed isotherm loops with continuously changing temperatures in the image, resulting in a set of closed isotherm loops. For each isotherm loop in the set, the server calculates the average temperature of the region inside the loop and the region immediately adjacent to the loop outside, obtaining internal and external average temperature parameters. The difference between these two parameters reflects the attenuation of thermal radiation when crossing the isotherm boundary and can be used to calculate a "temperature penetration boundary thickness" parameter, which indirectly characterizes the ability of thermal radiation to penetrate the smoke and dust medium. Similarly, through a mapping operation (penetration intensity mapping operation), the server uses this parameter to calculate a "radiation penetration coefficient" for each pixel in the thermal imaging image. A higher coefficient indicates that the thermal radiation at that point is less affected by smoke and dust and has better penetration. The radiation penetration coefficients of all pixels are arranged into a "radiation penetration feature map."
[0026] Next, the server inputs the obtained reflection attenuation feature map and radiation penetration feature map into a pre-trained "cross-modal feature reconstruction network." This network, deployed on the server, is designed to perform "feature compensation fusion operations." Specifically, the network learns to use the radiation penetration features of thermal imaging to compensate for and enhance details lost in visible light images in smoke-covered areas (such as personnel contour textures), generating an "enhanced personnel reflection feature map." Simultaneously, it uses environmental information provided by the visible light reflection attenuation features to calibrate and enhance the thermal radiation signal of personnel in thermal imaging, suppressing interference from non-personnel heat sources, generating an "enhanced personnel thermal radiation feature map." These two enhanced feature maps highlight the salient features of the personnel target in both modalities after smoke and dust environment compensation.
[0027] Then, the server performs a "stagnation personnel heat source tracking operation" based on these two enhanced feature maps. First, the server detects local temperature peaks in the enhanced personnel thermal radiation feature map and grows regions centered on these points to obtain a series of heat source connected domains. The temperature centroid of each connected domain is calculated as the "heat source spatial location coordinates." Simultaneously, the server detects closed curves of edge contours in the enhanced personnel reflection feature map and calculates the geometric center of each closed curve as the "reflection target spatial location coordinates." Subsequently, the server performs coordinate matching: for each heat source spatial location coordinate, it finds the closest Euclidean distance among the reflection target spatial location coordinates, and determines the pair of coordinates with the smallest distance and below a threshold as belonging to the same "candidate personnel target," forming a spatial location observation pair. The server repeats this process in consecutive multi-frame cross-modal image pairs, generating a time-spanning spatial location observation pair sequence for each candidate personnel target. To eliminate noise caused by detection jitter, the server filters this sequence (e.g., Kalman filtering) to remove distance anomalies, resulting in a smooth spatial location sequence. Finally, the server connects the coordinate points in this smooth sequence in chronological order of acquisition time to form a "heat source movement trajectory." The trajectories of all tracked candidate personnel constitute a "sequence set of heat source movement trajectories".
[0028] Finally, the server performs a "spatiotemporal dwelling pattern parsing operation" on the set of heat source movement trajectory sequences to identify lingering personnel. The server analyzes the morphology and spatiotemporal characteristics of each trajectory. First, it calculates the straight-line distance between the starting and ending points of the trajectory as the "overall trajectory displacement distance parameter." Simultaneously, it calculates the step length between adjacent coordinate points in the trajectory and takes the average of all step lengths as the "average movement step length parameter." Combining these two parameters, the server calculates a "trajectory curvature parameter" (e.g., the ratio of overall displacement to total path length, or estimated through curvature) to quantify the lingering characteristics of the trajectory. Trajectories with a curvature parameter exceeding a preset threshold (indicating that the target is moving back and forth in a local area rather than passing through in a straight line) are marked as "lingering motion trajectory sequences."
[0029] Next, the server performs a more refined temporal analysis on each loitering trajectory sequence. The server uses a sliding time window of fixed length to segment the trajectory sequence, resulting in multiple "trajectory sub-segments." For each sub-segment, the variance of the coordinates of all spatial coordinate points within it is calculated. Sub-segments with very small coordinate variances indicate that the target's spatial position changes minimally within that time period, placing it in a quasi-stationary state; therefore, they are marked as "resident sub-segments." The server records the start and end timestamps of each resident sub-segment. Subsequently, the server merges multiple resident sub-segments that are temporally adjacent and have a very short interval (e.g., less than 1 second) to form a "continuous resident period," and calculates the duration of this period as the "continuous resident duration parameter."
[0030] When a candidate's trajectory contains a period where the continuous dwell time exceeds a preset "dwell time threshold" (e.g., 30 seconds), the server marks the candidate as a "suspected lodger" and generates an identifier. The server extracts all spatial coordinates of the suspected lodger within the corresponding continuous dwell time period, calculates the average of these coordinates, and uses this as the "dwelling center location coordinate parameter." Finally, the server generates a structured "dwelling event record" for each suspected lodger, containing the person's identifier, the dwelling center location coordinates, and the continuous dwell time. The collection of all dwelling event records constitutes the system's output "tunnel lodger marking result," which can be directly used to trigger alarms or notify monitoring personnel. The entire process is executed automatically and continuously by the server, achieving intelligent detection of lodged behavior in tunnel environments.
[0031] In this embodiment of the invention, the complementary feature excitation operation performed on the cross-modal image frame pair to obtain the reflection attenuation feature map of the visible light image frame and the radiation penetration feature map of the thermal imaging image frame can be implemented through the following examples.
[0032] In this embodiment of the invention, for a visible light image frame, the server first converts it from the RGB color space to the grayscale space, obtaining a two-dimensional pixel brightness distribution matrix. Then, the server uses edge detection operators such as the Sobel operator to perform convolution operations on this matrix, calculating the brightness gradient components of each pixel in the horizontal and vertical directions, and then synthesizing the brightness attenuation gradient value (i.e., gradient magnitude) of that pixel. This value characterizes the degree of brightness change at that location; higher gradient values typically appear at the edges of smoke or dust or in areas of uneven density. The server arranges the brightness attenuation gradient values of all pixels according to their original coordinate positions, generating a brightness attenuation gradient distribution map of the same size as the original image. This map visualizes the distribution of regions with abrupt brightness changes in the image.
[0033] Next, the server performs gradient direction histogram statistics on this brightness attenuation gradient distribution map. The server divides the 360-degree gradient direction into several fixed intervals (e.g., every 20 degrees). Then, it iterates through each pixel in the distribution map, assigning it to the corresponding direction interval based on its gradient direction, and adding the gradient magnitude of that pixel to the total energy of that interval. After the statistics are completed, the server obtains a feature vector, where each dimension corresponds to the accumulated gradient energy of a direction interval; this is the gradient direction energy distribution feature vector. This vector reflects the main directional patterns of edges and textures in the image.
[0034] Based on this feature vector, the server applies a pre-trained smoke scattering region recognition model. This model, through learning, can identify regions with specific directional energy distribution patterns caused by suspended particulate matter. According to the model output, the server marks these smoke scattering regions on the original visible light image frame. Subsequently, the server extracts the brightness attenuation gradient values of all pixels within these regions obtained in previous calculations and calculates their arithmetic mean, which is defined as the smoke scattering intensity parameter for the current frame. This parameter globally characterizes the visual concentration of smoke within the tunnel at the current moment.
[0035] Finally, the server performs an attenuation mapping operation. Based on the smoke and dust scattering intensity parameters, it calculates a reflection attenuation coefficient for each pixel in the visible light image frame using a predefined mapping function (e.g., an exponential attenuation function). The input to this function is the pixel's original brightness value, and the output is a coefficient between 0 and 1; a lower coefficient indicates more severe attenuation of light from that point during propagation. The server arranges the reflection attenuation coefficients of all pixels by coordinates, ultimately generating a reflection attenuation feature map. This map numerically records the estimated attenuation of the visible light reflection signal at each point in the scene.
[0036] In the parallel processing thermal imaging branch, the server first reads the raw data of the thermal imaging image frame to obtain a two-dimensional pixel temperature value matrix. The server also uses a gradient operator to calculate the temperature gradient value of each pixel in the matrix, that is, a certain combination (such as the maximum value or root mean square value) of the temperature difference between the point and the temperature of its four or eight neighboring pixels, thereby generating a temperature gradient field distribution map.
[0037] Subsequently, the server performs isotherm tracing on the temperature gradient field distribution map. It employs contour-finding algorithms from image processing (such as topology-based algorithms) to identify closed curves, or isotherm loops, formed by pixels with continuous and similar temperature gradient values. Each loop represents the boundary of a relatively uniform temperature region. The server collects all identified loops, forming an isotherm loop set.
[0038] For each isotherm loop in the set, the server calculates the average temperature of all pixels inside the loop (internal average temperature parameter) and the average temperature of all pixels within a one-pixel-wide annular region immediately outside the loop (external average temperature parameter). The difference between these two parameters reflects the intensity of the temperature jump when crossing the isotherm boundary.
[0039] The server calculates the temperature penetration boundary thickness parameter of the isothermal ring using a formula inspired by a physical model, based on internal and external average temperature parameters. This parameter is inversely proportional to the intensity of the temperature jump; the more dramatic the jump, the "thinner" the boundary, indicating less resistance to heat radiation penetrating the boundary (possibly corresponding to the edge of a clearly visible heat source such as people); conversely, a gentler jump results in a "thicker" boundary, possibly corresponding to a heat diffusion area affected by smoke and dust dispersion.
[0040] Finally, the server performs a penetration intensity mapping operation. Based on the temperature penetration boundary thickness parameter of each isotherm loop, it assigns a radiation penetration coefficient to each pixel within the enclosed internal region and the adjacent external region. This coefficient is calculated using another mapping function; generally, the smaller the boundary thickness, the higher the coefficient, indicating better thermal radiation signal penetration and higher fidelity. The server traverses the entire thermal image, determining the isotherm loop to which each pixel belongs or its nearest neighbor, and calculates the corresponding radiation penetration coefficient, ultimately arranging these to generate a radiation penetration feature map. This map quantifies the relative strength of the thermal radiation signal's transmission capability at each point in the scene within a smoke and dust environment.
[0041] In this embodiment of the invention, the step of performing heat source tracking operation on stranded personnel based on the enhanced personnel thermal radiation feature map and the enhanced personnel reflection feature map to obtain a set of heat source movement trajectory sequences can be implemented through the following example.
[0042] Perform a local temperature peak detection operation on the enhanced personnel thermal radiation feature map to identify local peak pixels in the enhanced personnel thermal radiation feature map whose temperature values are greater than the temperature values of all pixels in the neighboring area, and obtain a set of local temperature peak points;
[0043] For each local temperature peak point in the set of local temperature peak points, a heat source region growth operation is performed. Using the local temperature peak point as a seed point, the operation expands and merges adjacent pixels whose temperature value decreases by a preset threshold to generate a heat source region connected component corresponding to each local temperature peak point.
[0044] Calculate the temperature centroid coordinates of each heat source region's connected domain, and use these temperature centroid coordinates as the spatial location coordinates of the heat source region's connected domain in the current cross-modal image frame pair;
[0045] An edge contour detection operation is performed on the enhanced human reflection feature map to identify the closed edge curves formed by connecting the local maxima of the gradient amplitude in the enhanced human reflection feature map, thereby obtaining a set of closed reflection contour curves.
[0046] Calculate the geometric center coordinates of the region enclosed by each reflection contour closed curve, and use the geometric center coordinates as the spatial position coordinates of the reflection target in the current cross-modal image frame pair;
[0047] The spatial coordinates of the heat source and the spatial coordinates of the reflecting target are input into the spatial position matching module. The Euclidean distance between each spatial coordinate of the heat source and each spatial coordinate of the reflecting target is calculated. The spatial coordinates of the heat source with the smallest Euclidean distance are paired with the spatial coordinates of the reflecting target as the spatial position observation pair of the same candidate human target.
[0048] Record the time series of spatial location observation pairs of each candidate target in a continuous cross-modal image frame pair sequence to obtain the original spatial location observation sequence of each candidate target;
[0049] An abnormal jump point filtering operation is performed on the original spatial location observation sequence to filter out abnormal observation points in the spatial location observation sequence whose distance value from the previous observation position exceeds a preset threshold, thereby generating a smooth spatial location sequence for each candidate target.
[0050] The smoothed spatial location sequence is connected sequentially according to the acquisition time sequence identifier to form a spatial coordinate change path, generating a heat source movement trajectory sequence for each candidate target, and combining the heat source movement trajectory sequences of all candidate targets into a heat source movement trajectory sequence set.
[0051] In an embodiment of the present invention, for example, after obtaining the enhanced thermal radiation feature map and the enhanced reflective feature map of the personnel, the server immediately performs a heat source tracking operation for the lingering personnel in order to construct the motion trajectory of each potential target.
[0052] First, the server processes the enhanced thermal radiation feature map of people. It uses a sliding window to traverse the map, comparing the temperature value of the center pixel with the values of all its neighboring pixels. When the center pixel's value is significantly greater than the values of all its neighbors, the server marks that pixel as a local temperature peak. After traversing the map, the server collects all such points, forming a set of local temperature peaks. These points represent potential independent heat source cores in the image.
[0053] Next, the server performs a heat source region growth operation on each local temperature peak point in the set. Using this peak point as a seed point, the server checks its neighboring pixels in four directions (up, down, left, right, and right) (or eight neighborhoods). If the descent gradient (difference) between the temperature value of a neighboring pixel and the seed point value is less than a preset gradient threshold, the server considers the neighboring points to belong to the same heat source region and merges them. Then, using the newly merged point as a new starting point, the server continues to perform the same judgment and merging process in all directions until there are no more neighboring pixels that meet the criteria. This process grows a connected pixel region for each peak point, i.e., the connected domain of the heat source region. Subsequently, the server calculates the temperature centroid coordinates of each connected domain of the heat source region. It takes a weighted average of the coordinates of each pixel within the connected domain and its temperature value, using the temperature value as the weight. The calculated weighted average coordinates are the precise spatial coordinates of the heat source in the current frame.
[0054] Meanwhile, the server processes the enhanced person reflection feature map in parallel. It uses the Canny edge detection algorithm or a similar operator to process this feature map. This algorithm first calculates the gradient magnitude and direction of the image, then performs non-maximum suppression to refine the edges, and finally identifies continuous pixel chains formed by connecting local maxima of gradient magnitude through double threshold detection and edge joining. The server analyzes these pixel chains, identifying those that form closed loops, thus forming a set of closed reflection contour curves. For each closed curve in the set, the server calculates the arithmetic mean of the coordinates of all the pixels it encloses to obtain the geometric center coordinates of the closed region. These coordinates are then recorded as the spatial location coordinates of a reflecting target in the current frame.
[0055] Subsequently, the server inputs the spatial coordinates of all heat sources and reflective targets in the current frame into the spatial location matching module. This module calculates the Euclidean distance between each heat source coordinate and each reflective target coordinate. For each heat source coordinate, the module finds the reflective target coordinate with the smallest Euclidean distance. If this minimum distance value is lower than a preset matching distance threshold, the server determines that this pair of coordinates belongs to the same physical entity (i.e., a candidate human target) and binds them into a "spatial location observation pair." This observation pair contains target location information jointly confirmed from both thermal radiation and visible light reflection modes.
[0056] When the server processes a sequence of consecutive multi-frame cross-modal image pairs, it maintains a list for each successfully tracked candidate target, recording in chronological order the spatial location observation pairs matched in each frame. This list constitutes the original spatial location observation sequence for that target.
[0057] Then, the server performs anomaly jump point filtering on the original sequence of each target. It iterates through each observation point in the sequence, starting from the second point, and calculates the Euclidean distance between that point and its previous observation point. If the calculated distance exceeds a reasonable threshold preset based on the average moving speed of the target, the server considers the current observation point to be an outlier (jump point) caused by mismatch or momentary occlusion, and removes it from the sequence. After filtering, the server obtains a smoothed spatial position sequence for each candidate target.
[0058] Finally, the server connects the coordinates of each target's smoothed spatial location sequence with line segments in a two-dimensional coordinate system, according to the chronological order of their corresponding acquisition time stamps. This connected path constitutes the heat source movement trajectory sequence of the candidate target. The server aggregates all tracked target trajectory sequences within the current monitoring period to form a heat source movement trajectory sequence set, providing a data foundation for subsequent dwell behavior analysis.
[0059] In this embodiment of the invention, the step of performing a spatiotemporal dwelling mode parsing operation on the set of heat source movement trajectory sequences to generate tunnel stranded personnel marking results can be implemented through the following example.
[0060] In this embodiment of the invention, for example, the server first iterates through each heat source movement trajectory sequence in the set. For each trajectory, the server extracts its first spatial coordinate point as the starting point and its last spatial coordinate point as the ending point. The server calculates the Euclidean distance between these two points, and this distance value is recorded as the "overall trajectory displacement distance parameter". This parameter reflects the net change in the target's spatial position over the entire observation period.
[0061] Next, the server calculates the Euclidean distance between two adjacent spatial coordinate points in the same trajectory sequence, i.e., the "step size value". The server arranges the step size values between all adjacent points in the trajectory in chronological order to form a "step size time series". Subsequently, the server calculates the arithmetic mean of all values in this time series to obtain the "average step size parameter", which characterizes the typical distance the target moves within a unit sampling interval.
[0062] Then, the server calculates a "trajectory curvature parameter" based on the "overall trajectory displacement distance parameter" and the "average movement step length parameter". This parameter is calculated by dividing the "overall trajectory displacement distance parameter" by the product of the "average movement step length parameter" and the "number of trajectory points minus one". The closer this ratio is to 0, the smaller the net displacement of the trajectory is compared to its cumulative movement path, indicating a high degree of trajectory curvature; the closer the ratio is to 1, the closer the trajectory is to linear motion. The server compares the calculated curvature parameter to a preset threshold (e.g., 0.3). Trajectory sequences with a curvature parameter below this threshold are labeled as "wandering motion trajectory sequences," considered to exhibit obvious local wandering characteristics and potential manifestations of lingering behavior.
[0063] For each marked wandering motion trajectory sequence, the server performs a spatiotemporal window segmentation operation. The server sets a fixed time window length (e.g., 10 seconds). Based on the timestamps of the trajectory points, the server divides the entire trajectory sequence into multiple consecutive "trajectory sub-segments" according to the time window, starting from the trajectory's start time. Each sub-segment contains all the spatial coordinate points collected within the corresponding time period, forming a "trajectory sub-segment set".
[0064] Subsequently, the server analyzes each trajectory sub-segment. It extracts the X and Y coordinates of all spatial coordinate points within the sub-segment, calculates the variance of the X and Y coordinates respectively, and then adds these two variance values to obtain the "coordinate variance value" of the sub-segment. The coordinate variance value quantifies the dispersion of the target's position within that time period. The server compares the coordinate variance value with a preset variance threshold. Sub-segments with coordinate variance values less than this threshold are marked as "stationary sub-segments," meaning that the target's position changes very little within this time period, and it is in a quasi-stationary state. The server records the start and end timestamps of each stationary sub-segment.
[0065] Next, the server merges temporally adjacent residency segments. It checks the temporal order of all residency segments; if the time interval between two segments (i.e., the difference between the end timestamp of the previous segment and the start timestamp of the next) is less than a preset merging interval threshold (e.g., 2 seconds), the server merges these two segments into a longer "continuous residency period." This process iterates until all temporally adjacent segments that meet the criteria have been merged. For each merged continuous residency period, the server calculates the difference between its start and end timestamps to obtain the "continuous residency duration parameter."
[0066] The server then compares the continuous dwell time parameter with a preset "dwell time threshold" (e.g., 30 seconds). For any period when the continuous dwell time parameter exceeds this threshold, the server traces back to the original heat source movement trajectory sequence that generated that period, marks the candidate personnel to which the trajectory sequence belongs as "suspected lingering personnel," assigns them a unique identifier, and adds them to the "suspected lingering personnel identifier list."
[0067] For each suspected stranded person in the list, the server extracts all spatial coordinate points within the "continuous stay period" of the trigger marker from the corresponding heat source movement trajectory sequence, forming a set of coordinate points. The server calculates the average of all X coordinates and the average of all Y coordinates in this set. These two averages together constitute the "stagnation center location coordinate parameters," which approximately represent the core area where the person is stranded.
[0068] Finally, the server generates a structured "Stallion Event Record" for each suspected stranded person. Each record contains at least three fields: the person's unique identifier (Stranded Person Identifier), the calculated coordinates of the stranded center location, and the continuous stay duration parameter that triggered the marker. The server aggregates all stranded event records generated within the current analysis period into a list or database. This aggregated result is the system's final output, the "Tunnel Stranded Person Marking Result," which can be used for subsequent alarm, visualization, or report generation.
[0069] In this embodiment of the invention, the method further includes performing a verification operation on the marking results of people stranded in the tunnel to generate a final confirmed set of detection results for people stranded in the tunnel, which can be implemented through the following example.
[0070] Extract all thermal radiation feature image frames of the enhanced thermal radiation feature image corresponding to each suspected stranded person in the tunnel stranded personnel marking results during the continuous residence period, and calculate the temperature-time curve of the temperature value at the coordinate parameter of the stranded center location in all thermal radiation feature image frames as a function of time.
[0071] Extract all reflection feature frames of the enhanced personnel reflection feature map corresponding to each suspected stranded person in the tunnel stranded personnel marking results during the continuous dwelling period, and calculate the reflection intensity time curve of the reflection intensity value at the coordinate parameter of the stranded center position in all reflection feature frames as a function of time.
[0072] Perform a first-order difference operation on the temperature-time curve to obtain a temperature change rate time series, and extract the frequency value of alternating positive and negative signs in the temperature change rate time series as the temperature fluctuation frequency parameter.
[0073] Perform a first-order difference operation on the reflection intensity time curve to obtain the reflection intensity change rate time series, and extract the frequency value of the alternating positive and negative signs in the reflection intensity change rate time series as the reflection intensity fluctuation frequency parameter;
[0074] The ratio of the temperature fluctuation frequency parameter to the reflection intensity fluctuation frequency parameter is calculated as the cross-modal fluctuation consistency coefficient.
[0075] The confidence score of the true stranded person is determined based on the cross-modal fluctuation consistency coefficient. When the cross-modal fluctuation consistency coefficient is within the preset consistency threshold range, the confidence score of the true stranded person is set as the first confidence value. When the cross-modal fluctuation consistency coefficient is outside the preset consistency threshold range, the confidence score of the true stranded person is set as the second confidence value.
[0076] Extract the preceding and following movement trajectory segments from the heat source movement trajectory sequence corresponding to each suspected stranded person in the tunnel stranded personnel marking results, outside the continuous dwelling period, and calculate the average movement direction angle of the preceding movement trajectory segment and the average movement direction angle of the following movement trajectory segment;
[0077] The absolute value of the difference between the average movement direction angle of the preceding movement trajectory segment and the average movement direction angle of the following movement trajectory segment is used as the consistency parameter of the inbound and outbound directions.
[0078] When the inbound / outbound direction consistency parameter is less than the preset direction consistency threshold, the inbound / outbound direction matching flag is set to the flag indicating that the person has entered but not left. When the inbound / outbound direction consistency parameter is greater than or equal to the preset direction consistency threshold, the inbound / outbound direction matching flag is set to the flag indicating that the person has entered but has left.
[0079] The confidence score of the true stranded person and the entry / exit direction matching identifier are input into the stranded person authenticity determination module. When the confidence score of the true stranded person is the first confidence value and the entry / exit direction matching identifier is the entry-not-leaved identifier, the suspected stranded person is confirmed as the true stranded person, and a true stranded person identifier is generated. The true stranded person identifier is combined with the corresponding stranded center location coordinate parameters, continuous stay duration parameters, and temperature values of the peak point of the temperature-time curve to form the final confirmed set of tunnel stranded person detection results.
[0080] In this embodiment of the invention, for example, firstly, for each suspected stranded person in the list, the server extracts key time-series data from its data cache. Specifically, based on the person's "stuck center location coordinate parameters," the server extracts the temperature value (or the average value of a small neighborhood) at that coordinate point from multiple consecutive frames of "enhanced person thermal radiation feature maps" within the corresponding continuous stay period, and arranges them in chronological order to form a "temperature-time curve." Simultaneously, the server extracts the reflection intensity value from multiple consecutive frames of "enhanced person reflection feature maps" at the same time period and coordinate location, and arranges them to form a "reflection intensity-time curve."
[0081] Next, the server performs rate-of-change analysis on both time curves. The server calculates the first-order difference sequence of the temperature-time curves (i.e., the difference in temperature values between adjacent time points) to obtain a temperature change rate sequence. The server counts the number of times the absolute value of the change rate in this sequence exceeds a small noise threshold, and divides this number by the total time length of the curve to obtain a "temperature fluctuation frequency parameter." This parameter quantifies the fluctuation activity of the target's surface thermal radiation during its stay. Similarly, the server performs the same processing on the reflection intensity time curve to calculate the "reflection intensity fluctuation frequency parameter," which quantifies the fluctuation activity of the target's surface reflection characteristics (potentially caused by minor movements or changes in clothing texture) during its stay.
[0082] The server then calculates the ratio of the "temperature fluctuation frequency parameter" to the "reflection intensity fluctuation frequency parameter" as the "cross-modal fluctuation consistency coefficient." A coefficient close to 1 indicates that the two physical signals exhibit similar fluctuation rhythms over time, which is usually a strong indication of subtle physiological activities or posture adjustments by a living person. The server predefines a reasonable coefficient range (e.g., [0.7, 1.3]). The server determines whether the calculated coefficient falls within this range: if it does, the target is considered to have high cross-modal fluctuation consistency, and a "high" level "true loiter confidence score" is assigned to it; if it falls outside the range, the confidence score is "low."
[0083] Simultaneously, the server performs directional consistency analysis. From the complete "heat source movement trajectory sequence" of the suspected lingering individual, it locates the starting point (the last movement point before entering the lingering state) and the ending point (the first movement point after leaving the lingering state) of the "continuous stay period." The server calculates the angle of movement from the point before the starting point to the starting point (e.g., using the arctangent function to calculate the angle relative to due east), denoted as the entry direction angle. Similarly, it calculates the angle of movement from the ending point to the point after the ending point, denoted as the exit direction angle. Then, the server calculates the absolute value of the difference between these two direction angles to obtain the "entry and exit direction consistency parameter." If this parameter value is less than a preset threshold (e.g., 90 degrees), it indicates that the target's entry and exit directions are roughly consistent, possibly indicating a brief stay followed by continued movement in the original direction; the server sets the "entry and exit direction matching identifier" to "left." If the parameter value is greater than or equal to the threshold, it indicates that the target's entry and exit directions differ significantly, possibly indicating lingering within the area before leaving in different directions; the server sets the "entry and exit direction matching identifier" to "not left" (suggesting a more likely lingering behavior).
[0084] Finally, the server performs a comprehensive decision. Only when a suspected stranded person's "True Stranded Person Confidence Score" is "High" and their "Entry / Exit Direction Matching Identifier" is "Not Left," is the server ultimately confirmed as a "True Stranded Person." For each confirmed true stranded person, the server generates an enhanced final record. This record not only includes the original identifier, the coordinates of the stranded center location, and the continuous stay duration, but also adds the highest peak value of the "Temperature-Time Curve" (temperature peak information) during the continuous stay period to reflect their thermal radiation intensity. The set of all final confirmed records constitutes the "Final Confirmed Tunnel Stranded Person Detection Result Set," which serves as the system's highest confidence output, used to trigger precise alarms or emergency responses.
[0085] In this embodiment of the invention, the method further includes performing a spatial location aggregation operation on the finally confirmed set of detection results of stranded personnel in the tunnel to generate a spatial distribution heat map of the stranded personnel, which can be implemented through the following example.
[0086] Extract the center position coordinates of each real stranded person from the final confirmed set of tunnel stranded personnel detection results, and map all center position coordinates to the two-dimensional plane coordinate system of the tunnel monitoring scene to obtain the distribution map of the center position coordinates.
[0087] A meshing operation is performed on the two-dimensional plane coordinate system to divide the two-dimensional plane coordinate system into multiple mesh cells of equal size. Each mesh cell has a unique mesh row index number and mesh column index number.
[0088] Traverse each stranded center coordinate point in the distribution map of stranded center coordinate points, determine the row index number and column index number of the grid cell to which each stranded center coordinate point belongs, and count the total number of stranded center coordinate points contained in each grid cell as the stranded personnel count value of that grid cell.
[0089] Extract the continuous dwell time parameters corresponding to all dwell center coordinates in each grid cell, and calculate the arithmetic mean of all continuous dwell time parameters in each grid cell as the average dwell time value of that grid cell;
[0090] Multiply the number of people staying in each grid cell by the average length of stay to obtain the cumulative weight of the stay in that grid cell. Calculate the maximum value of the cumulative weight of the stay in all grid cells as the normalized baseline value.
[0091] The normalized retention intensity coefficient of each grid cell is obtained by dividing the cumulative retention weight of each grid cell by the normalized reference value. The normalized retention intensity coefficient is then mapped to the corresponding color value in the preset color coding table to generate the grid fill color identifier for each grid cell.
[0092] Arrange the grid fill color identifiers of all grid cells into a color matrix according to the grid row index number and grid column index number, and perform bilinear interpolation amplification operation on the color matrix to generate a heat map of the spatial distribution of lingering personnel with continuous color tones.
[0093] The thermal map of the spatial distribution of stranded personnel is overlaid and fused with the visible light background image of the tunnel monitoring scene to generate a fused image of the tunnel monitoring scene with a thermal map of the spatial distribution of stranded personnel.
[0094] In this embodiment of the invention, for example, firstly, the server extracts the "detention center location coordinate parameters" contained in each real detainee record from the result set. These coordinate parameters are based on a two-dimensional plane coordinate system of the tunnel monitoring scenario (e.g., with the tunnel entrance as the origin, the X-axis along the tunnel direction, and the Y-axis perpendicular to the tunnel). The server plots all the extracted coordinate points in this coordinate system to form a "detention center coordinate point distribution map," which visually displays the spatial location of all detaining events in the form of points.
[0095] Next, the server performs a mesh generation operation on the aforementioned two-dimensional plane coordinate system. Based on the actual area of the tunnel monitoring scenario and the desired heatmap resolution, the server sets a fixed grid cell size (e.g., each grid represents 1 meter × 1 meter of actual space). Subsequently, the server uniformly divides the entire coordinate system from the minimum X value to the maximum X value and from the minimum Y value to the maximum Y value, generating a series of regularly arranged rectangular grid cells. Each grid cell is assigned a unique "grid row index number" and "grid column index number" for identification.
[0096] Then, the server iterates through each coordinate point in the "Distribution Map of Detention Center Coordinates". For each coordinate point, the server calculates and determines the row and column index of its corresponding grid cell based on its X and Y coordinate values. The server maintains a counter for each grid cell, incrementing the counter each time a coordinate point is determined to belong to that grid. After the iteration is complete, the final value of the counter for each grid cell is the "Detention Personnel Count" for that cell, indicating how many independent detention events have occurred within the physical area corresponding to that cell.
[0097] Subsequently, the server performs in-depth statistics. It iterates through the result set again, and for each real record of a stranded person, it not only determines the grid to which their coordinates belong but also extracts their "continuous stay duration parameter." The server sums the "continuous stay duration parameter" of all stranded events within each grid cell, and then divides it by the "stranded person count" of that grid cell to calculate the "average stay duration value" for that grid cell. This value reflects the average length of stay for stranded people in that area.
[0098] Next, the server calculates the "cumulative dwelling weight" for each grid cell. This value is calculated by multiplying the "dwelling personnel count" of that grid cell by its "average dwell time." This product takes into account both the frequency and severity (duration) of dwelling events; grid areas with higher values indicate more prominent dwelling problems. The server then finds the maximum value of the "cumulative dwelling weight" across all grid cells, using it as the "normalized baseline value."
[0099] For visualization mapping, the server divides the "cumulative retention weight" of each grid by the "normalized baseline value" to obtain a "normalized retention intensity coefficient" between 0 and 1. A coefficient of 0 indicates no retention, and a coefficient of 1 indicates the highest retention intensity. The server has a pre-defined color coding table (e.g., a gradient from blue [low intensity] to red [high intensity]). Based on the "normalized retention intensity coefficient" of each grid, the server looks up the corresponding RGB color value in the color table as the "grid fill color identifier" for that grid.
[0100] Then, the server arranges the "mesh fill color identifiers" of all grids into a two-dimensional "color matrix" according to the row and column order of the grid. Since the original grid may be relatively coarse, the server performs a "bilinear interpolation amplification operation" on the color matrix, smoothly inserting transition colors between grids, thereby generating a "spatial distribution heat map of stranded personnel" with continuous colors and smooth visual appearance. This heat map clearly shows the distribution of high and low risk of stranding in different areas of the tunnel.
[0101] Finally, the server merges the generated "spatial distribution heatmap of stranded personnel" with a clear "visible light background image" (as a base map) from the same tunnel section. The server sets the heatmap's transparency (e.g., 50%), and then, according to coordinate alignment, overlays the semi-transparent heatmap layer onto the visible light background image, generating a "fused image of the tunnel monitoring scene with a thermal distribution heatmap of stranded personnel." This fused image allows managers to clearly see the specific distribution of stranded events in the actual tunnel scene, greatly enhancing situational awareness.
[0102] In this embodiment of the invention, the method further includes performing a time-cumulative evolution operation on the marking results of people stranded in the tunnel to generate a time-cumulative evolution animation sequence of people stranded in the tunnel, which can be implemented through the following example.
[0103] Set the time window length parameter and the time window sliding step parameter of the sliding time window. Divide the entire monitoring period into multiple consecutive time window intervals according to the time window sliding step parameter. Each time window interval corresponds to a time window start timestamp and a time window end timestamp.
[0104] For each time window interval, perform a screening operation on the detection results of stranded personnel, extract the real stranded personnel in the final confirmed set of detection results of stranded personnel in the tunnel whose continuous stay period overlaps with the current time window interval, and generate a window stranded personnel subset corresponding to each time window interval;
[0105] Based on the coordinates of the center of the actual stranded persons in each subset of stranded persons in each window, a spatial location aggregation operation is performed to generate a heat map of the spatial distribution of stranded persons in each time window interval.
[0106] All heatmaps showing the spatial distribution of people staying at all windows are arranged into a heatmap time series according to the temporal order of the time window intervals. Pixel value difference operations are performed on two adjacent heatmaps in the heatmap time series to obtain a heatmap change difference map series.
[0107] The sum of the absolute values of the pixel values of each heatmap change difference map in the heatmap change difference map sequence is extracted as the heatmap change intensity parameter. Heatmap change difference maps whose heatmap change intensity parameter exceeds the preset change intensity threshold are marked as key change frame identifiers.
[0108] Based on the key change frame identifier, extract the window lingering personnel spatial distribution heat map corresponding to the key change frame from the heat map time series, and perform image overlay and fusion operation on the window lingering personnel spatial distribution heat map corresponding to the key change frame and the visible light background image of the tunnel monitoring scene to generate the fused map of lingering personnel spatial distribution at the key change moment.
[0109] The spatial distribution map of stranded personnel at all key change moments is arranged into a key frame image sequence in chronological order. Image frame interpolation is performed between adjacent frames in the key frame image sequence to generate an interpolated transition frame image set.
[0110] The keyframe image sequence and the interpolated transition frame image set are merged into a continuous animation frame sequence in chronological order, and the animation frame sequence is encoded into a time-accumulated lingering personnel evolution animation sequence according to a preset playback frame rate.
[0111] In this embodiment of the invention, for example, the server first sets two key parameters: the time window length (e.g., 5 minutes) and the time window sliding step size (e.g., 1 minute). Starting from the beginning of the entire monitoring period, the server sequentially extracts consecutive time window intervals at intervals equal to the sliding step size. Each interval is defined by its start and end timestamps, and adjacent intervals overlap in time. The server divides the entire monitoring period into a series of such time window intervals.
[0112] Next, the server performs a filtering operation on the detection results of stranded personnel for each time window interval. It iterates through each real stranded personnel record in the "final confirmed set of tunnel stranded personnel detection results". The server checks whether the "continuous stay period" (defined by its start and end timestamps) in the record overlaps with the current time window interval. If the two time periods overlap, the server determines that the stranding event occurred within the current time window and adds the record to the "window stranded personnel subset" corresponding to the current window. This operation filters out stranded personnel who are active within that time period for each time window.
[0113] Then, for each "window-dwelling personnel subset," the server performs a spatial location aggregation operation similar to that described in claim 6. Specifically, the server extracts the "dwelling center location coordinate parameters" of all records in the subset and maps them to the tunnel's two-dimensional plane coordinate system. Next, the server uses the same grid division rules to count the number of dwelling center coordinate points appearing in each grid cell (i.e., the dwelling frequency within that window) and calculates the average value of the "continuous dwelling duration parameter" corresponding to these points. Subsequently, the server calculates the cumulative dwelling weight value (frequency × average duration) for each grid, normalizes it to obtain the "normalized dwelling intensity coefficient," and maps it to color. Finally, the server generates an independent "window-dwelling personnel spatial distribution heatmap" for each time window. This map reflects the spatial distribution of dwelling intensity within the tunnel during that specific time segment.
[0114] Subsequently, the server arranges the heatmaps for all time windows in chronological order, forming a "heatmap time series." The server analyzes this series frame by frame: calculating the difference in RGB color values of corresponding pixels in the color space between two adjacent heatmaps (i.e., frame n and frame (n+1)), summing the absolute values of the differences for each pixel to obtain a "heatmap change difference map," the sum of which is the "heatmap change intensity parameter." This parameter quantifies the degree of change in the persistence distribution between two adjacent time windows. The server marks the next heatmap frame corresponding to a difference map where the change intensity parameter exceeds a preset threshold (e.g., 10% of the total pixel difference) as a "critical change frame." This identifies the moment when the persistence distribution undergoes a significant change.
[0115] Subsequently, based on the "key change frame identifiers," the server extracts the corresponding "spatial distribution heatmaps of stranded personnel" from the "heatmap time series." The server then semi-transparently overlays and fuses each keyframe heatmap with the "visible light background image" of the tunnel monitoring scene, generating a series of "fused spatial distribution maps of stranded personnel at key change moments." These fused maps visually demonstrate the location of stranded hotspots in the actual tunnel scene at the points in time when significant changes occurred in the stranding situation.
[0116] Finally, the server performs animation sequence synthesis. It arranges all the "spatial distribution fusion maps of stranded personnel at key change moments" in chronological order to form a "keyframe image sequence." Since keyframes may be discontinuous in time (intervals of several minutes), to achieve a smooth animation effect, the server performs image frame interpolation operations between adjacent keyframes (e.g., using motion estimation and compensation algorithms, or simple fade-in / fade-out and deformation transitions), generating a series of "interpolated transition frame images." The server interweaves and merges the original keyframes and the generated transition frames in the correct chronological order to form a "continuous animation frame sequence" with a uniform frame rate (e.g., 10 frames per second). Finally, the server uses a video encoder (such as H.264) to compress and encode this animation frame sequence, generating the final "time-cumulative evolution animation sequence of stranded personnel." This animation can be played dynamically, clearly showing how the distribution of stranded personnel in the tunnel evolves, gathers, and dissipates over time throughout the monitoring period.
[0117] In this embodiment of the invention, the method further includes performing an abnormal loitering pattern recognition operation on the set of heat source movement trajectory sequences to generate an abnormal loitering event warning signal, which can be implemented through the following example.
[0118] Extract the start and end timestamps of all heat source movement trajectory sequences in the heat source movement trajectory sequence set, determine the earliest start and latest end timestamps of the entire monitoring period, and divide the time period between the earliest start timestamp and the latest end timestamp into multiple equal-length intervals.
[0119] The number of active heat source movement trajectory sequences in the set of heat source movement trajectory sequences within each equal time interval is counted as the active heat source count value for that time interval. The active heat source count values of all time intervals are arranged in chronological order to form an active heat source count time curve.
[0120] Perform local peak detection on the active heat source counting time curve to identify local peak points and local valley points in the active heat source counting time curve, and calculate the active heat source count change amplitude between adjacent local peak points and local valley points;
[0121] The active heat source count rise segment where the change value of the active heat source count exceeds the preset amplitude threshold is marked as the personnel gathering event segment. The start time stamp and end time stamp of the personnel gathering event segment are extracted, and the difference between the start time stamp and end time stamp of the personnel gathering event segment is calculated as the personnel gathering duration parameter.
[0122] When the duration of the gathering exceeds the preset gathering duration threshold, the spatial coordinates of all heat source movement trajectory sequences corresponding to the gathering event segment are extracted within the gathering event segment, and the coordinate variance of all spatial coordinates is calculated as the gathering spatial dispersion parameter.
[0123] When the dispersion parameter of the gathering space is less than the preset dispersion threshold, the personnel gathering event segment is marked as an abnormal dense gathering event, and a dense gathering event timestamp identifier and dense gathering space center coordinate parameters are generated.
[0124] Extract the trajectory curvature parameter of all heat source movement trajectory sequences in the heat source movement trajectory sequence set, mark the heat source movement trajectory sequences whose trajectory curvature parameter exceeds the preset hovering threshold as hovering trajectory sequences, and count the number of hovering trajectory sequences in each equal time interval as the hovering trajectory count value of that time interval;
[0125] The ratio of the loitering trajectory count to the active heat source count over multiple consecutive time intervals is used as the loitering personnel ratio parameter. When the loitering personnel ratio parameter exceeds the preset ratio threshold and the number of consecutive time intervals exceeds the preset number threshold, the corresponding time period is marked as a continuous growth event of loitering personnel, and the start timestamp and end timestamp of the growth event are generated.
[0126] The abnormally dense gathering events and the continuous increase in the number of stranded people are input into the early warning signal generation module to generate an early warning signal for abnormal stranding events, which includes an abnormal event type identifier, an abnormal event occurrence timestamp, and abnormal event spatial location coordinates.
[0127] In this embodiment of the invention, for example, firstly, the server extracts the start and end timestamps of all trajectories from the set of heat source movement trajectory sequences. The server finds the minimum value among these timestamps as the "earliest start timestamp" for the entire monitoring period and the maximum value as the "latest end timestamp". The server divides this complete time period into multiple consecutive, fixed-length "equal-length intervals" (e.g., each interval is 2 minutes long). These intervals serve as the basic units for subsequent time series analysis.
[0128] Next, the server counts the number of active heat source movement trajectory sequences within each equal time interval, obtaining the "active heat source count value" for that interval. The server determines whether a trajectory is active within a certain time interval based on the following rule: at least one spatial coordinate point of the trajectory has a timestamp falling within this time interval. The server arranges the count values of all time intervals in chronological order, forming an "active heat source count time curve," which reflects the change in the number of tracked targets within the tunnel over time.
[0129] Subsequently, the server performs local peak detection on this active heat source count time curve. It uses a sliding window or signal processing algorithm to identify local peaks (points with relatively high counts) and local troughs (points with relatively low counts) on the curve. For each pair of adjacent "local trough -> local peak," the server calculates the difference between the peak and trough counts to obtain the "active heat source count change amplitude value." This amplitude value quantifies the intensity of the increase in the number of people over a short period.
[0130] The server marks any increase in the "active heat source count change value" exceeding a preset threshold (e.g., an increase of more than 10 targets in a short period of time) as a "personnel gathering event segment". The server records the start timestamp (valley point time) and end timestamp (peak point time) of this event segment and calculates the difference between the two as the "personnel gathering duration parameter". When this duration parameter exceeds a preset threshold (e.g., the gathering lasts for more than 1 minute), the server considers it a persistent gathering of concern.
[0131] Then, the server performs spatial analysis on the persistent clustering event. It extracts all spatial coordinate points generated within the "personnel clustering event segment" time period, based on the movement trajectory sequences of all active heat sources. The server calculates the variance of these coordinate points in the X and Y directions and sums them to obtain the "clustering spatial dispersion parameter." If this parameter is less than a preset dispersion threshold, it indicates that these active targets are spatially highly concentrated, forming an abnormally dense cluster, and the server marks this event as an "abnormally dense clustering event." The server records the "dense clustering event timestamp" (usually the peak time) and the "dense clustering spatial center coordinate parameter" (the mean of all coordinate points within that time period).
[0132] On another parallel analysis line, the server utilizes the previously calculated "trajectory curvature parameter" for each trajectory. The server sets a high "wandering threshold," marking trajectories with curvature parameters exceeding this threshold as "wandering trajectory sequences." The server counts the number of active "wandering trajectory sequences" within each "equal time interval," obtaining the "wandering trajectory count value" for that interval.
[0133] The server calculates a key metric: for each equal-length interval, it divides the "loitering trajectory count" of that interval by the "active heat source count" to obtain the "loitering personnel percentage parameter." This parameter reflects the proportion of people exhibiting loitering behavior during the current time period. The server monitors the time series of this percentage parameter. When it detects that the "loitering personnel percentage parameter" exceeds a preset percentage threshold (e.g., 50%) for multiple consecutive (e.g., 5 consecutive) equal-length intervals, the server marks this consecutive time period as a "continuous growth event of loitering personnel." This indicates that not only is there loitering within the tunnel, but the loitering trend is also spreading or intensifying. The server records the "growth event start timestamp" and "growth event end timestamp" of this growth event.
[0134] Finally, the server inputs the identified "abnormally dense gathering events" and "continuous increase in stranded personnel events" into the early warning signal generation module. This module generates a structured "abnormal stranding event early warning signal" for each abnormal event. Each signal contains at least three core fields: an abnormal event type identifier (such as "dense gathering" or "trend growth"), an abnormal event occurrence timestamp (corresponding to the peak or center time of the growth phase), and an abnormal event spatial location coordinate parameter (gathering center or center of the growth area). These early warning signals are pushed to the monitoring center in real time, providing managers with high-level alerts regarding the risk of group stranding.
[0135] In this embodiment of the invention, the method further includes performing a graded response operation on the early warning signal of the abnormal lodging event to generate a set of emergency response instructions at different levels, which can be implemented through the following example.
[0136] Analyze the abnormal event type identifier in the abnormal lodging event warning signal. When the abnormal event type identifier is an abnormal dense clustering event, extract the spatial center coordinate parameters of the dense clustering event and the timestamp identifier of the dense clustering event.
[0137] The straight-line distance between the center coordinates of the densely clustered space and the location coordinates of the tunnel safety exit is calculated as the distance parameter between the cluster point and the exit. The distance parameter between the cluster point and the exit is then compared with preset near-exit distance thresholds and far-exit distance thresholds.
[0138] When the distance parameter between the gathering point and the exit is less than the near exit distance threshold, a Level 1 emergency response instruction is generated. The Level 1 emergency response instruction includes an exit evacuation guidance sign and a gathering point broadcast notification code.
[0139] When the distance parameter between the gathering point and the exit is greater than or equal to the near exit distance threshold and less than the far exit distance threshold, a level 2 emergency response instruction is generated. The level 2 emergency response instruction includes an area isolation identifier and a patrol personnel dispatch code.
[0140] When the distance parameter between the gathering point and the exit is greater than or equal to the far exit distance threshold, a Level 3 emergency response instruction is generated. The Level 3 emergency response instruction includes a video surveillance focus indicator and a remote announcement activation code.
[0141] The abnormal event type identifier in the early warning signal of abnormal stay events is analyzed. When the abnormal event type identifier is a continuous increase in the number of stayers, the start time stamp and end time stamp of the growth event are extracted, and the difference between the start time stamp and the end time stamp of the growth event is calculated as the duration parameter of the growth event.
[0142] The average growth rate parameter of the growth event is calculated based on the duration parameter of the growth event, and then compared with the preset high-speed growth rate threshold and low-speed growth rate threshold.
[0143] When the average growth rate parameter is greater than the high-speed growth rate threshold, a first-level growth response instruction is generated. The first-level growth response instruction includes a tunnel entrance flow restriction indicator and a reinforcement personnel dispatch code.
[0144] When the average growth rate parameter is less than or equal to the high-speed growth rate threshold and greater than the low-speed growth rate threshold, a secondary growth response instruction is generated. The secondary growth response instruction includes a tunnel broadcast reminder sign and a key focus code for the monitoring center.
[0145] When the average growth rate parameter is less than or equal to the low growth rate threshold, a level 3 growth response instruction is generated. The level 3 growth response instruction includes a regular patrol enhancement identifier and a historical data record archiving code. All levels of emergency response instructions are arranged into an emergency response instruction sequence according to the generation time, and the emergency response instruction sequence is encapsulated into a transmissible emergency response instruction set.
[0146] In this embodiment of the invention, for example, firstly, the server parses the "abnormal event type identifier" in the warning signal. When the identifier is "abnormal dense clustering event," the server extracts the "dense clustering spatial center coordinate parameters" and "dense clustering event timestamp identifier" attached to the signal. The server internally stores tunnel plan data, which presets the location coordinates of all safety exits. The server calculates the Euclidean distance between the clustering center coordinates and the coordinates of the nearest safety exit to obtain the "clustering point and exit distance parameter." The server presets two distance thresholds: "near exit distance threshold" (e.g., 20 meters) and "far exit distance threshold" (e.g., 100 meters).
[0147] The server compares this distance parameter with these two thresholds. If the distance is less than the "near-exit distance threshold," it indicates that the dense gathering is occurring near the exit, which is highly likely to cause congestion and stampedes, posing the highest risk. The server then generates a "Level 1 Emergency Response Instruction." This instruction contains two core operational codes: one is an "exit evacuation guidance sign," which triggers the corresponding exit's indicator sign to flash and switches to emergency evacuation mode; the other is a "gathering point broadcast notification code," which controls directional broadcasting equipment near the gathering point to play evacuation prompts.
[0148] If the distance parameter is greater than or equal to the "near exit distance threshold" but less than the "far exit distance threshold," it indicates that the clustering occurred in the middle area of the tunnel. The server generates a "Level 2 Emergency Response Command." This command includes an "Area Isolation Identifier," which sends a descent or activation command to the nearest automatic isolation device (such as a roller shutter or warning post) to softly isolate the clustering area; it also includes a "Patrol Personnel Dispatch Code," which pushes the clustering location and event information to the nearest patrol personnel's handheld terminal, dispatching them to handle the situation.
[0149] If the distance parameter is greater than or equal to the "distant exit distance threshold", it indicates that the aggregation is occurring deep within the tunnel, far from the exit. The server generates a "Level 3 Emergency Response Command". This command includes a "Video Surveillance Focus Identifier", which automatically controls the nearest high-definition PTZ camera to turn and focus on the coordinates of the aggregation center for detailed monitoring; it also includes a "Remote Broadcast Activation Code", which activates the remote broadcast system equipped in the area, allowing monitoring center personnel to intervene via voice.
[0150] When the server analyzes the warning signal and finds that the "abnormal event type identifier" is "continuous increase in stranded personnel event," it adopts a different set of evaluation logic. The server extracts the "growth event start timestamp" and "growth event end timestamp" from the signal and calculates the difference between them to obtain the "growth event duration parameter." The server divides the total increase in active heat source counts during this period by the duration to obtain the "average growth rate parameter" (unit: people / minute). The server presets a "high-speed growth rate threshold" (e.g., 5 people / minute) and a "low-speed growth rate threshold" (e.g., 2 people / minute).
[0151] The server compares the average growth rate parameter with these two thresholds. If the parameter is greater than the "high-speed growth rate threshold," it indicates that the trend of people remaining in the tunnel is deteriorating rapidly, and the risk is escalating quickly. The server generates a "Level 1 Growth Response Instruction," which includes a "tunnel entrance flow restriction sign," sending instructions to the entrance gate or signal light system to temporarily restrict new personnel from entering; it also includes a "reinforcement personnel dispatch code," notifying the emergency response team to provide reinforcements.
[0152] If the average growth rate parameter is less than or equal to the high-speed threshold but greater than the low-speed threshold, it indicates a clear but still controllable growth trend. The server generates a "Level 2 Growth Response Command," which includes a "tunnel broadcast reminder flag" to activate a global broadcast within the tunnel, playing a safety reminder; it also includes a "key focus code for the monitoring center," highlighting the growth area on the monitoring screen to remind the duty officer to pay close attention.
[0153] If the average growth rate parameter is less than or equal to the "slow growth rate threshold", it indicates a slow growth trend. The server generates a "Level 3 Growth Response Instruction", which includes a "Regular Patrol Enhancement Identifier" to prompt patrol personnel to increase patrol frequency in the area in the future; it also includes a "Historical Data Record Archiving Code" to archive the detailed parameters of the event for subsequent analysis.
[0154] Finally, the server arranges all emergency response instructions generated according to the above rules at all levels in chronological order of their generation, forming an ordered "emergency response instruction sequence." The server encapsulates this sequence into a structured data packet, namely the "emergency response instruction set," and sends it in real time to various terminal execution devices (such as broadcasting, turnstiles, and cameras) and personnel dispatching systems within the tunnel via a dedicated communication link, thus completing the closed loop from intelligent analysis to automatic response.
[0155] In this embodiment of the invention, the method further includes performing a bidirectional verification update operation on the enhanced human thermal radiation feature map and the enhanced human reflectance feature map to generate an adaptively optimized cross-modal feature reconstruction network parameter set, which can be implemented through the following example.
[0156] The enhanced personnel thermal radiation feature map is compared with the corresponding original thermal imaging image frame in the original thermal imaging image frame sequence to perform pixel-by-pixel difference calculation to obtain the thermal radiation feature reconstruction error distribution map. The sum of squares of the error values of all pixels in the thermal radiation feature reconstruction error distribution map is calculated as the thermal radiation feature reconstruction loss parameter.
[0157] The enhanced personnel reflection feature map is compared with the corresponding original visible light image frame in the original visible light image frame sequence to perform pixel-by-pixel difference calculation to obtain the reflection feature reconstruction error distribution map. The sum of squares of the error values of all pixels in the reflection feature reconstruction error distribution map is calculated as the reflection feature reconstruction loss parameter.
[0158] The locations of pixels whose error values exceed a preset error threshold in the thermal radiation feature reconstruction error distribution map are extracted as the set of thermal radiation reconstruction defect points, and the locations of pixels whose error values exceed a preset error threshold in the reflection feature reconstruction error distribution map are extracted as the set of reflection reconstruction defect points.
[0159] Calculate the number of intersection pixels between the set of defect points reconstructed by thermal radiation and the set of defect points reconstructed by reflection, and divide the number of intersection pixels by the total number of pixels in the set of defect points reconstructed by thermal radiation to obtain the cross-modal defect overlap ratio parameter.
[0160] When the cross-modal defect overlap ratio parameter exceeds the preset overlap ratio threshold, the pixel position of the union of the thermal radiation reconstruction defect point set and the reflection reconstruction defect point set is marked as the cross-modal common defect region, and a common defect region mask map is generated.
[0161] Based on the common defect region mask map, extract the original visible light image patch and the original thermal imaging image patch corresponding to the common defect region from the original cross-modal image frame pair, and use the original visible light image patch and the original thermal imaging image patch as incremental training samples to input the cross-modal feature reconstruction network;
[0162] In the cross-modal feature reconstruction network, backpropagation parameter update operation is performed. The gradient of the loss function of the cross-modal feature reconstruction network is calculated based on the incremental training samples. The convolution kernel weight parameters and bias parameters of the cross-modal feature reconstruction network are adjusted in the opposite direction of the loss function gradient.
[0163] Record the thermal radiation feature reconstruction loss parameters and reflection feature reconstruction loss parameters corresponding to each parameter update operation. When the mean of the thermal radiation feature reconstruction loss parameters and reflection feature reconstruction loss parameters decreases less than the preset convergence threshold after multiple consecutive parameter update operations, stop the parameter update operation.
[0164] Save the convolutional kernel weights and biases of the cross-modal feature reconstruction network when the parameter update operation is stopped as an adaptively optimized set of parameters for the cross-modal feature reconstruction network.
[0165] In an embodiment of the invention, for example, the server performs a two-way verification update operation to optimize the core network. This operation first calculates the pixel-by-pixel error between the enhanced feature map and the corresponding original image, obtaining the reconstruction error distribution maps of thermal radiation and reflection features and the corresponding overall loss parameters.
[0166] The server analyzes the error distribution and extracts defective pixels exceeding a preset threshold from two error maps, forming two reconstructed defect point sets: thermal radiation and reflection. Then, it calculates the number of pixels overlapping between these two sets and divides it by the total number of thermal radiation defect points to obtain a "cross-modal defect overlap ratio parameter." When this ratio exceeds a preset threshold, it indicates the existence of a consistent performance defect region across modes. The server merges the two defect point sets to generate a "common defect region mask map" that identifies these difficult regions.
[0167] Next, the server enters the incremental training phase. It extracts image patches of corresponding regions from the original image pairs based on the mask image, using these patches as incremental training samples input to the "cross-modal feature reconstruction network." The network calculates the loss gradient through forward and backward propagation, and fine-tunes all its weights and bias parameters accordingly to reduce reconstruction errors in these challenging regions.
[0168] The server continuously monitors the updated loss parameters. When the average decrease in the loss parameters is less than a preset convergence threshold after the most recent parameter updates, the server stops updating. At this point, the server saves the final parameter state of the network completely, forming an "adaptive optimization cross-modal feature reconstruction network parameter set". This parameter set can be used to update the network, achieving continuous self-optimization of the system.
[0169] In this embodiment of the invention, the method further includes performing a spatiotemporal oscillation excitation operation on cross-modal image frame pairs to generate an oscillation enhancement feature map of a visible light image frame for personnel detection and a thermal inertia response feature map of a thermal imaging image frame. This can be implemented through the following examples.
[0170] In an embodiment of the present invention, for example, in addition to the basic processing flow, the server also executes an auxiliary analysis flow called "spatiotemporal oscillation excitation operation" in parallel, which aims to mine deep features of image sequences from the time dimension and generate oscillation enhancement feature maps and thermal inertia response feature maps for assisting personnel detection.
[0171] First, the server retrieves cross-modal image frame pairs from the cache, including the current moment and several consecutive moments prior (e.g., the most recent 10 frames), and separates the visible light image frames to form a "visible light image frame sequence." For each pixel in the sequence (determined by its coordinates [x, y]), the server extracts its brightness value in each frame along the time axis, forming a "pixel brightness time series." The server performs a second-order difference operation on this time series, that is, calculates the first-order difference (velocity) between adjacent brightness values, and then calculates the difference on the first-order difference sequence to obtain a "pixel brightness acceleration sequence," which reflects the acceleration of the brightness change at that point.
[0172] Next, the server analyzes the brightness acceleration sequence of each pixel. It identifies points in the sequence where the sign of the values changes; these points are called "abrupt change points," indicating a shift in brightness change from acceleration to deceleration or vice versa. The server records the "acquisition timing identifier" of the image frame corresponding to each abrupt change point, serving as a "brightness oscillation inflection point" for that pixel. Then, the server counts the total number of times the brightness oscillation inflection point for that pixel occurs within the most recent fixed unit of time (e.g., 1 second), defining this number as the "pixel brightness oscillation frequency value." A higher frequency value indicates more drastic fluctuations in brightness within a short period.
[0173] Subsequently, the server arranges the calculated "pixel brightness oscillation frequency values" of all pixels according to their original two-dimensional coordinate positions, generating a "brightness oscillation frequency distribution map" of the same size as the original image. This map visually displays the frequency of brightness fluctuations at each point in the entire scene over time in grayscale values. The server performs a connected component analysis algorithm on this distribution map, grouping pixels with similar frequency values and spatially adjacent locations together to form several "brightness oscillation synchronization regions," resulting in a "brightness oscillation synchronization region set."
[0174] The server then analyzes each region in the set. It calculates the area of each region (the number of pixels it contains) and the average oscillation frequency value of all pixels within that region (the mean oscillation frequency parameter). Regions where the mean oscillation frequency parameter exceeds a preset threshold (e.g., more than 5 fluctuations per second) are marked as "regions with drastically fluctuating smoke concentration." This is because suspended smoke particles cause rapid, random fluctuations in scattered light intensity, exhibiting temporal characteristics different from a static background or a steadily moving object. The server generates a binary "smoke fluctuation region mask" based on the coordinates of these regions.
[0175] Finally, the server uses this mask image to enhance the original visible light image frame. It extracts all pixels marked as smoke fluctuation regions by the mask image from the original visible light image of the current frame. The server then multiplies the original brightness values of these pixels by a "time decay factor" (e.g., 0.8) less than 1 to simulate the average brightness decay caused by the time integration effect of rapid fluctuations, obtaining "time decay compensated brightness values." The server fills these compensated brightness values back into the corresponding positions in the original image, while the brightness of pixels in unmarked areas remains unchanged, ultimately generating an "oscillation enhancement feature map of the visible light image frame." This map suppresses brightness flicker noise caused by rapid smoke fluctuations, improving the temporal stability of the image.
[0176] In the parallel processing branch of thermal imaging, the server performs a similar but physically different operation. It extracts multiple consecutive "thermal imaging image frame sequences" within the same time window. For each pixel, the server arranges its temperature values in each frame in chronological order, forming a "pixel temperature time series." The server then performs numerical integration on this time series (e.g., using the trapezoidal rule to calculate the cumulative sum) to obtain a "pixel temperature accumulation curve," which reflects the cumulative process of heat absorption or release at that point over a period of time.
[0177] Next, the server calculates the area under the cumulative curve within a unit of time (e.g., the most recent second), defining this area value as the "pixel thermal inertia response value". Points with high thermal inertia response values indicate that their temperature changes significantly over time, potentially corresponding to living individuals (whose body temperature is relatively stable compared to the environment, creating a continuous contrast with the background during movement or slight motion) or continuous heat sources. The server arranges the thermal inertia response values of all pixels by coordinates to generate an "initial thermal inertia response distribution map".
[0178] In this embodiment of the invention, the method further includes performing a dwell phase locking operation on the oscillation enhancement feature map and the thermal inertia response feature map to generate a phase-locked dwell personnel detection result, which can be implemented through the following example.
[0179] In this embodiment of the invention, for example, the server first performs pixel-level alignment of two feature maps. For each identical pixel coordinate, the server takes the "brightness oscillation frequency value" in the oscillation enhancement feature map and the "thermal inertia response value" in the thermal inertia response feature map as a single observation sample of two time series and inputs it into the "phase synchronization determination module." This module calculates the normalized cross-correlation function between these two values. Since the input is an instantaneous value rather than a long sequence, this calculation actually assesses the potential correlation strength between the changes in the two signals on a more macroscopic time scale (the frequency and response values calculated from previous frames already contain time information). The server extracts the theoretical "delay time parameter" corresponding to when the cross-correlation function reaches its maximum value as the "cross-modal response delay time" of that pixel. A very small delay time means that the brightness fluctuation and thermal inertia change at that point occur almost synchronously.
[0180] Next, the server marks all pixels with a "cross-modal response delay time" less than a preset threshold (e.g., 0.1 seconds) as "phase-synchronous pixels." The server generates a binary image, namely a "phase-synchronous pixel binary mask," where synchronized pixels are white (value 1) and others are black (value 0). Then, the server performs a morphological closing operation (dilation followed by erosion) on the mask to fill the breaks between synchronized pixels caused by noise or tiny gaps, thereby obtaining connected regions and forming a "phase-synchronous connected region set."
[0181] The server then performs geometric shape filtering on each connected region. It calculates the area (total number of pixels) and perimeter of each region's outline. The server uses the ratio of area to perimeter as a "region compactness parameter" (similar to a roundness index; higher compactness indicates a shape closer to a circle or dense cluster). The server filters out regions with a compactness parameter below a preset threshold (e.g., elongated strips or sparsely distributed regions), deeming them inconsistent with the shape of typical human targets. The remaining regions constitute the "filtered phase-synchronized region set."
[0182] Then, the server extracts features from each of the filtered regions. Based on the region coordinates, it extracts corresponding image patches from the oscillation enhancement feature map and the thermal inertia response feature map, respectively. The server calculates the average of all pixel feature values within each image patch to obtain the region's "regional oscillation frequency feature" and "regional thermal inertia response feature." The server combines these two feature values into a two-dimensional "phase-locked intensity feature vector."
[0183] The feature vector is input into a pre-trained "dwelling phase classifier" (e.g., a support vector machine or a small neural network). Based on historical data, the classifier learns that there is a specific strong correlation between the brightness oscillations (potentially reflecting micro-motions) and thermal inertia response (reflecting sustained thermal characteristics) in the area where a person is truly in a dwelling state. The classifier outputs a "dwelling phase lock probability value" between 0 and 1.
[0184] The server marks areas where the probability value exceeds a preset locking probability threshold (e.g., 0.7) as "phase-locked personnel areas". The server calculates the centroid coordinates of each marked area as "phase-locked personnel spatial location coordinates" and records the "acquisition timing identifier" of the current processing frame as "phase-locking time parameter".
[0185] During continuous frame processing, the server performs a "phase-locked personnel tracking operation." It compares the coordinates of each phase-locked personnel detected in the current frame with the coordinates detected in the previous frame. If the Euclidean distance between two coordinates is less than a "tracking distance threshold" set based on the maximum possible movement speed, the server associates them as the same "phase-locked personnel target" and updates the target's trajectory sequence.
[0186] Finally, the server analyzes the trajectory sequence of each target. It calculates the magnitude (scale) of the displacement vector between adjacent time points in the trajectory. The server marks a continuous period of time where the displacement magnitude is consistently less than a "preset static displacement threshold" (e.g., 0.5 meters) as a "phase-locked dwell time period" and records the start and end times of this period. The duration of this period is calculated to obtain the "phase-locked dwell time parameter". When this duration exceeds a preset "phase-locked dwell time threshold" (e.g., 20 seconds), the server ultimately classifies the target as a "phase-locked loitering person". The server generates a structured detection result for such persons, including a unique identifier, their spatial location coordinate sequence within the dwell time period, and the dwell time, forming an independent "phase-locked loitering person detection result" that can be cross-validated with the main process results.
[0187] In this embodiment of the invention, the method further includes performing adversarial game enhancement operations on the reflection attenuation feature map and the radiation penetration feature map to generate adversarial game optimized reflection attenuation feature map and adversarial game optimized radiation penetration feature map, which can be implemented through the following examples.
[0188] The reflection attenuation feature map is input into the first feature discriminator network. The first feature discriminator network performs a convolution downsampling operation on the reflection attenuation feature map to extract the global statistical distribution parameters of the reflection attenuation feature map. Based on the global statistical distribution parameters, the probability value of the reflection attenuation feature map belonging to the true reflection attenuation feature is calculated as the true probability score of the reflection feature.
[0189] The radiation penetration feature map is input into the second feature discriminator network. The second feature discriminator network performs a convolution downsampling operation on the radiation penetration feature map to extract the global statistical distribution parameters of the radiation penetration feature map. Based on the global statistical distribution parameters, the probability value of the radiation penetration feature map belonging to the true radiation penetration feature is calculated as the true probability score of the radiation feature.
[0190] The reflection attenuation feature map and the radiation penetration feature map are input into the cross-modal feature perturbation generator. The cross-modal feature perturbation generator performs deformable convolution operation on the reflection attenuation feature map to generate a reflection feature perturbation field, and performs deformable convolution operation on the radiation penetration feature map to generate a radiation feature perturbation field.
[0191] The perturbation field of the reflection feature is superimposed on the reflection attenuation feature map to generate the perturbation reflection attenuation feature map, and the perturbation field of the radiation feature is superimposed on the radiation penetration feature map to generate the perturbation radiation penetration feature map. The perturbation reflection attenuation feature map is input into the first feature discriminator network to calculate the true probability score of the perturbation reflection feature, and the perturbation radiation penetration feature map is input into the second feature discriminator network to calculate the true probability score of the perturbation radiation feature.
[0192] The difference between the true probability score of the reflection feature and the true probability score of the reflection feature after perturbation is calculated as the reflection feature discrimination loss value, and the difference between the true probability score of the radiation feature and the true probability score of the radiation feature after perturbation is calculated as the radiation feature discrimination loss value.
[0193] When the reflection feature discrimination loss value is less than the preset reflection discrimination loss threshold, the current parameters of the cross-modal feature perturbation generator are fixed. When the reflection feature discrimination loss value is greater than or equal to the preset reflection discrimination loss threshold, the gradient of the reflection feature discrimination loss value is backpropagated to update the parameters of the cross-modal feature perturbation generator.
[0194] When the radiation feature discrimination loss value is less than the preset radiation discrimination loss threshold, the current parameters of the cross-modal feature perturbation generator are fixed. When the radiation feature discrimination loss value is greater than or equal to the preset radiation discrimination loss threshold, the gradient of the radiation feature discrimination loss value is backpropagated to update the parameters of the cross-modal feature perturbation generator.
[0195] A pixel-by-pixel additive fusion operation is performed between the reflection attenuation feature map and the final reflection feature perturbation field generated by the parameter-updated cross-modal feature perturbation generator to generate an adversarial game-optimized reflection attenuation feature map. A pixel-by-pixel additive fusion operation is performed between the radiation penetration feature map and the final radiation feature perturbation field generated by the parameter-updated cross-modal feature perturbation generator to generate an adversarial game-optimized radiation penetration feature map.
[0196] In this embodiment of the invention, for example, the server performs an "adversarial game enhancement operation" on the generated reflection attenuation feature map and radiation penetration feature map to optimize their feature representation. The server first inputs the two feature maps into independent feature discriminator networks. Each discriminator evaluates the realism of the input feature map in its statistical distribution and outputs a true probability score. Then, the server inputs the original feature map into a cross-modal feature perturbation generator. This generator learns to generate a corresponding pixel-level perturbation field for each feature map. The server superimposes the perturbation field onto the original feature map to generate a perturbed feature map, which is then fed back into the discriminator for evaluation to obtain a new true probability score. Next, the server calculates the difference between the original score and the perturbed score as a discriminant loss value. This loss value measures the impact of the perturbation on the "realism" of the feature. Based on this loss value, the server updates the parameters of the perturbation generator using a backpropagation algorithm, aiming to minimize the discriminant loss, i.e., to ensure that the perturbed feature map still appears "realistic" or even more so to the discriminator. After iterative game-based training, the server uses an optimized perturbation generator to produce the final perturbation field and fuses it pixel-by-pixel with the original feature map, ultimately generating adversarial game-optimized reflection attenuation feature maps and radiation penetration feature maps. This process enhances the robustness and discriminative power of the features.
[0197] In this embodiment of the invention, the method further includes performing a reverse feature decoupling operation on the adversarial game-optimized reflection attenuation feature map and the adversarial game-optimized radiation penetration feature map to generate a reflection attenuation source component feature map and a radiation penetration source component feature map, which can be implemented through the following example.
[0198] The reflection attenuation feature map optimized by adversarial game is input into the reflection feature encoding branch of the feature decoupling encoder. The reflection feature encoding branch performs a depthwise separable convolution operation on the reflection attenuation feature map optimized by adversarial game to extract the channel attention weight vector of the reflection attenuation feature. Based on the channel attention weight vector, a weighting and calibration operation is performed on each feature channel of the reflection attenuation feature map optimized by adversarial game to generate a reflection attenuation weighted feature tensor.
[0199] The adversarial game-optimized radiation penetration feature map is input into the radiation feature encoding branch of the feature decoupling encoder. The radiation feature encoding branch performs a depthwise separable convolution operation on the adversarial game-optimized radiation penetration feature map to extract the channel attention weight vector of the radiation penetration feature. Based on the channel attention weight vector, a weighting and calibration operation is performed on each feature channel of the adversarial game-optimized radiation penetration feature map to generate a radiation penetration weighted feature tensor.
[0200] The reflection attenuation weighted feature tensor and the radiation penetration weighted feature tensor are input into the joint feature decoupling module of the feature decoupling encoder. The joint feature decoupling module calculates the covariance matrix of the reflection attenuation weighted feature tensor and the radiation penetration weighted feature tensor, and performs eigenvalue decomposition on the covariance matrix to obtain the covariance eigenvalues and covariance eigenvectors.
[0201] The covariance eigenvector corresponding to the largest eigenvalue among the covariance eigenvalues is extracted as the main correlation direction vector. The reflection attenuation weighted feature tensor is projected onto the main correlation direction vector to obtain the reflection attenuation main correlation component. The radiation penetration weighted feature tensor is projected onto the main correlation direction vector to obtain the radiation penetration main correlation component.
[0202] The reflection attenuation residual component is obtained by subtracting the reflection attenuation principal correlation component from the reflection attenuation weighted feature tensor. The radiation penetration residual component is obtained by subtracting the radiation penetration principal correlation component from the radiation penetration weighted feature tensor. The reflection attenuation residual component is input into the reflection source component decoder. The reflection source component decoder performs a transpose convolution operation on the reflection attenuation residual component to generate the reflection attenuation source component feature map.
[0203] The radiation penetration residual component is input into the radiation source component decoder. The radiation source component decoder performs a transpose convolution operation on the radiation penetration residual component to generate a radiation penetration source component feature map. The reflection attenuation main correlation component and the radiation penetration main correlation component are input into the cross-modal shared component decoder. The cross-modal shared component decoder performs a feature fusion convolution operation on the reflection attenuation main correlation component and the radiation penetration main correlation component to generate a cross-modal shared interference component feature map.
[0204] The structural similarity index between the reflection attenuation source component feature map and the reflection attenuation feature map optimized by adversarial game is calculated as the reflection source component retention parameter, and the structural similarity index between the radiation penetration source component feature map and the radiation penetration feature map optimized by adversarial game is calculated as the radiation source component retention parameter.
[0205] When the retention parameter of the reflection source component is less than the preset retention threshold, the gradient of the retention parameter of the reflection source component is backpropagated to update the convolution kernel parameter of the reflection source component decoder. When the retention parameter of the radiation source component is less than the preset retention threshold, the gradient of the retention parameter of the radiation source component is backpropagated to update the convolution kernel parameter of the radiation source component decoder.
[0206] The final reflection attenuation source component feature map output by the parameter-updated reflection source component decoder is used as the final output of the reverse feature decoupling operation, and the final radiation penetration source component feature map output by the parameter-updated radiation source component decoder is used as the final output of the reverse feature decoupling operation.
[0207] In this embodiment of the invention, for example, firstly, the server inputs the "adversarial game-optimized reflection attenuation feature map" into the "reflection feature encoding branch" of a "feature decoupling encoder". This branch is a lightweight neural network that processes the input feature map using "depthiable separable convolution". This convolution performs spatial convolution and pointwise convolution respectively, efficiently extracting features. The last convolutional layer of the branch outputs a "channel attention weight vector", the dimension of which is equal to the number of channels in the input feature map, and each value represents the importance of the corresponding feature channel. The server uses this weight vector to scale (i.e., perform weighting and calibration) each channel of the original optimized feature map, enhancing channels with high weights and suppressing those with low weights, thereby generating a "reflection attenuation weighted feature tensor" that focuses on key information.
[0208] At the same time, the server inputs the "radiation penetration feature map optimized for adversarial game" into the "radiation feature encoding branch" of the same encoder. This branch has the same structure but independent parameters, performs the exact same operation, and finally generates a "radiation penetration weighted feature tensor".
[0209] Next, the server inputs these two weighted feature tensors into the core of the encoder, the "joint feature decoupling module." This module first flattens the two tensors in spatial dimensions into a set of feature vectors, and then calculates the "covariance matrix" between these two sets to measure the degree of linear correlation between the two modal features. The server performs an "eigenvalue decomposition" operation on this covariance matrix to obtain a set of eigenvalues and corresponding eigenvectors.
[0210] The server extracts the "covariance eigenvector" corresponding to the largest eigenvalue from the decomposition results. This vector is called the "main correlation direction vector". It represents the strongest common change pattern between the two modal features, which usually corresponds to shared interference factors such as smoke and dust interference and changes in ambient light.
[0211] Subsequently, the server performs a projection operation to separate the components. It projects the "reflection attenuation weighted feature tensor" along the "main correlation direction vector" to obtain the "reflection attenuation main correlation component," which contains the part of the reflection feature that is highly correlated with the thermal radiation feature (i.e., shared interference). Similarly, the "radiation penetration weighted feature tensor" is projected to obtain the "radiation penetration main correlation component." Then, the server performs a subtraction operation: subtracting the "reflection attenuation main correlation component" from the original "reflection attenuation weighted feature tensor" to obtain the "reflection attenuation residual component"; and subtracting the "radiation penetration main correlation component" from the "radiation penetration weighted feature tensor" to obtain the "radiation penetration residual component." These residual components theoretically contain the mode-specific feature information after removing the main cross-modal shared interference.
[0212] Next, the server inputs the "reflection attenuation residual component" into a dedicated "reflection source component decoder." This decoder, composed of multiple layers of transposed convolutions, is responsible for upsampling the abstract representation of the residual component and reconstructing it back into an image of the same size as the original feature map, generating a "reflection attenuation source component feature map." This map aims to preserve and enhance the essential reflection features related to people in the visible light modality. Similarly, the server processes the "radiation penetration residual component" through a "radiation source component decoder" to generate a "radiation penetration source component feature map."
[0213] In addition, the server inputs the extracted "reflection attenuation main correlation component" and "radiation penetration main correlation component" into a "cross-modal shared component decoder". This decoder fuses and upsamples these two components to generate a "cross-modal shared interference component feature map", which visualizes the separated common noise patterns.
[0214] To ensure the effectiveness of the decoupling operation, the server performs fidelity verification. It calculates the Structural Similarity Index (SSIM) between the generated "Reflection Attenuation Source Component Feature Map" and the original "Adversarial Game-Optimized Reflection Attenuation Feature Map," obtaining a "Reflection Source Component Retention Parameter" between 0 and 1. This parameter measures the degree to which the decoupled feature map retains structural information from the original map. The server presets a "Retention Threshold." If the calculated parameter is less than this threshold, it indicates that the decoder reconstruction loss is too large. The server then uses the gradient of this parameter to update the convolution kernel parameters of the "Reflection Source Component Decoder" through backpropagation to improve its reconstruction capability. The exact same verification and parameter update process is performed for the thermal radiation mode.
[0215] After iterative optimization, the server uses the feature map output by the finally stable "reflection source component decoder" as the "final output of the reverse feature decoupling operation," i.e., the pure "reflection attenuation source component feature map." Similarly, the feature map output by the finally stable "radiation source component decoder" is used as the "radiation penetration source component feature map." These two source component feature maps, which have eliminated shared interference, will be used for subsequent more robust and accurate feature fusion and personnel analysis.
[0216] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned method and system for detecting people stranded in tunnels. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0217] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A method for detecting people stranded in a tunnel, characterized in that, The method includes: The visible light image frame sequence and the thermal imaging image frame sequence of the tunnel section are acquired. Each visible light image frame in the visible light image frame sequence and the thermal imaging image frame sequence with the same acquisition time identifier constitute a cross-modal image frame pair. Complementary feature excitation operation is performed on the cross-modal image frame pair to obtain the reflection attenuation feature map of the visible light image frame and the radiation penetration feature map of the thermal imaging image frame. The reflection attenuation feature map records the propagation loss distribution of visible light in the tunnel smoke environment, and the radiation penetration feature map records the transmission intensity distribution of thermal radiation in the tunnel smoke environment. The reflection attenuation feature map and the radiation penetration feature map are input into a cross-modal feature reconstruction network to perform feature compensation and fusion operations, generating an enhanced personnel thermal radiation feature map and an enhanced personnel reflection feature map; Based on the enhanced thermal radiation feature map and the enhanced reflection feature map of the personnel, a heat source tracking operation is performed on the stranded personnel to obtain a set of heat source movement trajectory sequences. The set of heat source movement trajectory sequences includes the spatial coordinate change path of each heat source in a continuous cross-modal image frame pair. A spatiotemporal dwelling pattern parsing operation is performed on the set of heat source movement trajectory sequences to generate a tunnel dwelling personnel marking result, which includes a dwelling start time parameter and a dwelling duration parameter.
2. The method according to claim 1, characterized in that, The step of performing complementary feature excitation on the cross-modal image frame pair to obtain the reflection attenuation feature map of the visible light image frame and the radiation penetration feature map of the thermal imaging image frame includes: Extract the pixel brightness distribution matrix of the visible light image frame in the cross-modal image frame pair, calculate the brightness attenuation gradient value of each pixel in the visible light image frame based on the pixel brightness distribution matrix, and arrange the brightness attenuation gradient values according to the pixel coordinate position to obtain the brightness attenuation gradient distribution map of the visible light image frame. Perform gradient direction histogram statistics on the brightness attenuation gradient distribution map, and calculate the cumulative sum of gradient magnitudes in each gradient direction interval to generate gradient direction energy distribution feature vector of visible light image frame; Based on the gradient direction energy distribution feature vector, the smoke and dust scattering region in the visible light image frame is identified, and the mean value of the brightness attenuation gradient of the pixels in the smoke and dust scattering region is extracted as the smoke and dust scattering intensity parameter. The smoke scattering intensity parameter is mapped to the original brightness value of each pixel in the visible light image frame to generate a reflection attenuation coefficient for each pixel. The reflection attenuation coefficients are then arranged into a reflection attenuation feature map according to the pixel coordinates. Extract the pixel temperature value matrix of the thermal imaging image frame in the cross-modal image frame pair, calculate the difference between the temperature value of each pixel and the temperature value of its neighboring pixels in the pixel temperature value matrix as the temperature gradient value, and generate the temperature gradient field distribution map of the thermal imaging image frame. Perform isotherm tracking operation on the temperature gradient field distribution map to identify isotherm closed loops in the thermal imaging image frame where the temperature gradient value changes continuously, and obtain a set of isotherm closed loops. Calculate the average internal temperature value and the average external temperature value of each isotherm closed loop in the isotherm closed loop set to obtain the internal average temperature parameter and the external average temperature parameter. The temperature penetration boundary thickness parameter of each isotherm closed loop is calculated based on the internal average temperature parameter and the external average temperature parameter. The temperature penetration boundary thickness parameter is then mapped to the original temperature value of each pixel in the thermal imaging image frame to generate the radiation penetration coefficient corresponding to each pixel. The radiation penetration coefficients are then arranged into a radiation penetration feature map according to the pixel coordinate position.
3. The method according to claim 1, characterized in that, The step of performing heat source tracking operation on stranded personnel based on the enhanced personnel thermal radiation feature map and the enhanced personnel reflection feature map yields a set of heat source movement trajectory sequences, including: Local temperature peaks in the enhanced personnel thermal radiation feature map are identified, and a region is grown with each peak as the center to obtain a connected domain of the heat source region. The temperature centroid of each connected domain is calculated as the spatial coordinate of the heat source. Identify the closed curves of the edge contours in the enhanced human reflection feature map, and calculate the geometric center of each closed curve as the spatial position coordinates of the reflection target; The spatial coordinates of the heat source are matched with the spatial coordinates of the reflecting target, and the pair of coordinates with the smallest Euclidean distance is determined as the spatial position observation pair of the same candidate personnel target; Obtain the spatial location observation pair sequence of each candidate target in consecutive image frame pairs, filter out distance outliers in the sequence, and obtain a smooth spatial location sequence; The smooth spatial position sequence is connected in chronological order to form a spatial coordinate change path, which serves as the heat source movement trajectory. The set of all heat source movement trajectories constitutes the heat source movement trajectory sequence set.
4. The method according to claim 1, characterized in that, The step of performing a spatiotemporal dwell pattern parsing operation on the set of heat source movement trajectory sequences to generate tunnel stranded personnel marking results includes: Extract the starting and ending spatial coordinates of each heat source movement trajectory sequence from the heat source movement trajectory sequence set, and calculate the straight-line distance between the starting and ending spatial coordinates as the overall trajectory displacement distance parameter; Calculate the step size between adjacent spatial coordinates in each heat source movement trajectory sequence, arrange all step size values in chronological order into a step size time series, and calculate the arithmetic mean of all step size values in the step size time series as the average step size parameter. The trajectory curvature parameter of each heat source movement trajectory sequence is calculated based on the overall trajectory displacement distance parameter and the average movement step length parameter. Heat source movement trajectory sequences whose trajectory curvature parameter exceeds a preset threshold are marked as wandering motion trajectory sequences. Perform a spatiotemporal window segmentation operation on each wandering motion trajectory sequence, dividing the wandering motion trajectory sequence into multiple trajectory sub-segments according to a fixed time window length, and obtain a set of trajectory sub-segments; Calculate the coordinate variance of spatial coordinate points in each trajectory sub-segment, mark the trajectory sub-segments with coordinate variance values less than a preset threshold as stationary sub-segments, and record the start and end timestamps of each stationary sub-segment; Multiple time-adjacent dwell segments with a time interval less than a preset threshold are merged into a continuous dwell period, and the difference between the start and end timestamps of the continuous dwell period is calculated as the continuous dwell duration parameter. The candidate personnel whose continuous dwell time parameter exceeds the preset dwell time threshold are marked as suspected stranded personnel, and a list of suspected stranded personnel identifiers is generated. Extract the set of spatial coordinate points of the heat source movement trajectory sequence corresponding to each suspected stranded person during the continuous residence period, and calculate the mean of the coordinates of the set of spatial coordinate points as the coordinate parameters of the stranded center location; Based on the location coordinates of the detention center and the duration of continuous stay, a detention event record is generated for each suspected detainee. The detention event record includes the detainee identifier, the location coordinates of the detention center, and the duration of continuous stay. All records of stranded events are compiled to generate a result of marking stranded personnel in the tunnel.
5. The method according to claim 4, characterized in that, The method further includes: Extract the temperature-time curve and reflection intensity-time curve at the center of the stay for each suspected stranded person during the continuous stay period; Calculate the rate of change sequences of the temperature-time curve and the reflection intensity-time curve, and obtain the temperature fluctuation frequency parameter and the reflection intensity fluctuation frequency parameter, respectively; The ratio of the temperature fluctuation frequency parameter to the reflection intensity fluctuation frequency parameter is calculated and used as the cross-modal fluctuation consistency coefficient. The confidence score of the true stranded person is determined based on whether the cross-modal fluctuation consistency coefficient is within a preset range. Take the movement direction angle of the trajectory of the suspected stranded person before and after the continuous stay period, and calculate the absolute value of the difference in direction angle as the consistency parameter of the inbound and outbound directions. Based on the comparison result between the inbound / outbound direction consistency parameter and the preset threshold, an inbound / outbound direction matching identifier is determined; When the confidence score of the true stranded person is high and the entry / exit direction matching identifier indicates that the person has not left, the suspected stranded person is confirmed as a true stranded person, and a final confirmed set of tunnel stranded person detection results containing identifier, location, duration and temperature peak information is generated.
6. The method according to claim 5, characterized in that, The method further includes: Map the coordinates of all the locations of the stranded centers in the detection result set to the tunnel's two-dimensional plane coordinate system; divide the two-dimensional plane coordinate system into a grid, and count the number of stranded people and the average stay time in each grid; The cumulative value of the dwelling weight is calculated by multiplying the number of people staying in each grid by the average length of stay, and then normalized to obtain the normalized dwelling intensity coefficient. The normalized retention intensity coefficient is mapped to a color value, and the corresponding grid is filled to generate a grid color matrix. The grid color matrix is interpolated and enlarged to generate a continuous spatial distribution heat map of lingering personnel; The spatial distribution heatmap of the stranded personnel is fused with the visible light background image of the tunnel to generate a fused image of the tunnel monitoring scene with a heatmap layer.
7. The method according to claim 5, characterized in that, The method further includes: The entire monitoring period is divided into multiple consecutive time window intervals by using a sliding time window. For each time window interval, filter out the actual stranded persons whose stay time overlaps with that time period, and generate a subset of stranded persons within the window; Based on the location information of each subset of people staying at the window, a corresponding spatial distribution heat map of people staying at the window is generated; Arrange all window heatmaps in chronological order to obtain a heatmap time series, and calculate the pixel differences between adjacent heatmaps; Heatmaps with differences exceeding a preset threshold are marked as key change frames and fused with the tunnel background image to generate a fused map of key moments; The fused images of the key moments are arranged in sequence to form a key frame sequence, and transition frames are inserted between adjacent key frames to form a continuous animation frame sequence. The animation frame sequence is encoded according to a preset frame rate to generate the time-accumulated lingering personnel evolution animation sequence.
8. The method according to claim 1, characterized in that, The method further includes the step of performing a spatiotemporal oscillation excitation operation on cross-modal image frame pairs to generate an oscillation enhancement feature map for visible light image frames and a thermal inertia response feature map for thermal imaging image frames for personnel detection. This step includes: Extract the visible light image frame sequence of multiple consecutive cross-modal image frame pairs from the cross-modal image frame pair, arrange the brightness value of each pixel in the visible light image frame sequence in the time dimension into a pixel brightness time sequence, and perform a second-order difference operation on the pixel brightness time sequence to obtain a pixel brightness acceleration sequence. Identify the locations of abrupt changes in the positive and negative signs of pixel brightness acceleration sequences, record the acquisition timing identifiers corresponding to the abrupt changes as brightness oscillation inflection points, and count the number of times each pixel's brightness oscillation inflection points occur per unit time as pixel brightness oscillation frequency values. Arrange the pixel brightness oscillation frequency values of all pixels according to their pixel coordinates to form a brightness oscillation frequency distribution map of the visible light image frame. Perform connected component analysis on the brightness oscillation frequency distribution map to extract regions with continuous brightness oscillation frequency values, and obtain a set of brightness oscillation synchronization regions. Calculate the area parameter and mean oscillation frequency parameter of each brightness oscillation synchronization region in the brightness oscillation synchronization region set. Mark the brightness oscillation synchronization regions whose mean oscillation frequency parameter exceeds the preset oscillation frequency threshold as regions with violent fluctuations in smoke and dust concentration, and generate a smoke and dust fluctuation region mask map. Based on the smoke and dust fluctuation area mask map, the set of pixels corresponding to the area of violent smoke and dust concentration fluctuation is extracted from the original visible light image frame. The original brightness value of the set of pixels corresponding to the area of violent smoke and dust concentration fluctuation is multiplied by the time decay factor to obtain the time decay compensation brightness value. The time decay compensation brightness value is filled back into the area of violent smoke and dust concentration fluctuation to obtain the oscillation enhancement feature map of the visible light image frame. Extract thermal imaging image frame sequences from multiple consecutive cross-modal image frame pairs, arrange the temperature values of each pixel in the thermal imaging image frame sequence in the time dimension into a pixel temperature time series, and perform an integral operation on the pixel temperature time series to obtain the pixel temperature cumulative curve. The area under the cumulative temperature curve of a pixel per unit time is calculated as the pixel thermal inertia response value. The pixel thermal inertia response values of all pixels are arranged according to the pixel coordinate position to form the initial distribution map of thermal inertia response of the thermal imaging image frame. Gaussian smoothing is performed on the initial distribution map of thermal inertia response to eliminate pixel-level noise and obtain a smooth distribution map of thermal inertia response. Local maxima of thermal inertia response values in the smooth distribution map of thermal inertia response are extracted as peak points of thermal inertia response. The difference between the peak point of thermal inertia response and the thermal inertia response values of its surrounding pixels is taken as the magnitude of thermal inertia response gradient. Based on the magnitude of the thermal inertia response gradient, a region growth segmentation operation is performed on the thermal inertia response smooth distribution map. Starting from each thermal inertia response peak point, the operation expands outwards and merges adjacent pixels whose thermal inertia response gradient magnitude is less than a preset gradient threshold, generating a set of connected regions of the thermal inertia response. Calculate the centroid coordinates and total response parameters of each connected domain in the set of thermal inertia response regions. Mark the connected domains of thermal inertia response regions whose total response parameters exceed the preset total response threshold as personnel thermal inertia response regions, and generate thermal inertia response feature maps of thermal imaging image frames.
9. The method according to claim 8, characterized in that, The method further includes the step of performing a dwell phase-locking operation on the oscillation enhancement feature map and the thermal inertia response feature map to generate a phase-locked dwell personnel detection result, which includes: The brightness oscillation frequency value of each pixel in the oscillation enhancement feature map and the thermal inertia response value of the pixel with the same coordinates in the thermal inertia response feature map are input into the phase synchronization determination module. The cross-correlation function between the brightness oscillation frequency value and the thermal inertia response value is calculated, and the delay time parameter corresponding to the maximum value of the cross-correlation function is extracted as the cross-modal response delay time of the pixel. Pixels with cross-modal response delay times less than a preset delay time threshold are marked as phase-synchronous pixels. A phase-synchronous pixel binary mask is generated. A morphological closing operation is performed on the phase-synchronous pixel binary mask to fill the tiny gaps between the phase-synchronous pixels, resulting in a set of phase-synchronous connected regions. Calculate the area parameter and perimeter parameter of each phase-synchronized connected region in the set of phase-synchronized connected regions. Use the ratio of the area parameter to the perimeter parameter as the region compactness parameter. Filter out phase-synchronized connected regions whose region compactness parameter is less than a preset compactness threshold to obtain the filtered set of phase-synchronized regions. Extract the oscillation enhancement feature map region block and thermal inertia response feature map region block corresponding to each phase synchronization region in the filtered phase synchronization region set. Calculate the average value of the brightness oscillation frequency of all pixels in the oscillation enhancement feature map region block as the region oscillation frequency feature, and calculate the average value of the thermal inertia response of all pixels in the thermal inertia response feature map region block as the region thermal inertia response feature. A phase-locking strength feature vector is constructed based on the regional oscillation frequency characteristics and the regional thermal inertia response characteristics. The phase-locking strength feature vector is then input into a pre-trained resident phase classifier, which outputs the resident phase-locking probability value of the phase synchronization region. The phase synchronization region where the dwell phase lock probability value exceeds the preset lock probability threshold is marked as the phase lock personnel region. The region centroid coordinates of each phase lock personnel region are extracted as the spatial position coordinates of the phase lock personnel. The acquisition time sequence identifier of the current cross-modal image frame pair is recorded as the phase lock time parameter. For multiple consecutive cross-modal image frame pairs, perform phase-locked personnel tracking operation, associate phase-locked personnel regions in adjacent cross-modal image frame pairs where the distance between the spatial coordinates of phase-locked personnel is less than a preset tracking distance threshold as the same phase-locked personnel target, and generate a spatiotemporal trajectory sequence of the phase-locked personnel target; Calculate the displacement vector between the spatial coordinates of the phase-locked personnel at adjacent moments in the spatiotemporal trajectory sequence of the phase-locked personnel target. Mark the continuous time segments where the magnitude of the displacement vector is less than the preset static displacement threshold as the phase-locked dwell time period. Record the start and end time parameters of the phase-locked dwell time period. The difference between the start time parameter and the end time parameter of the phase-locked dwell period is calculated as the phase-locked dwell duration parameter. When the phase-locked dwell duration parameter exceeds the preset phase-locked dwell duration threshold, the phase-locked personnel target is marked as a phase-locked stranded personnel, and a phase-locked stranded personnel detection result containing the phase-locked personnel identifier, the phase-locked personnel spatial location coordinate sequence, and the phase-locked dwell duration parameter is generated.
10. A system for detecting people stranded in a tunnel, characterized in that, It includes at least one service node; the service node includes a storage unit and a computing unit; the storage unit is used to store program code; the computing unit is used to run the program code to execute the tunnel personnel detection method according to any one of claims 1 to 9.