A composite abnormal behavior recognition method and system in a fire safety and security coupled scene
By mapping two-dimensional image coordinates to a top-down physical coordinate system, a scene geometric semantic model is constructed and dynamically updated, solving the problems of inconsistent spatial measurements and misjudgment of behavior recognition in fire protection and security coupled scenarios, and realizing effective identification of cross-scenario adaptability and multi-dimensional risk warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING KUNYA TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies lack a unified physical scale benchmark, dynamic topology reconstruction mechanism, and multi-element coupling calculation mechanism in fire protection and security coupled scenarios, resulting in inconsistent spatial measurements, misaligned benchmarks for judging reverse behavior, and insufficient group risk linkage early warning.
By mapping two-dimensional image coordinates to a top-down physical coordinate system based on the calibration parameters of the surveillance camera, a scene geometric semantic model is constructed, and abandoned object detection and contour extraction are performed. The obstacle interference analysis is dynamically updated, and evacuation behavior is identified and reverse behavior detection results are generated.
It enables cross-scenario adaptive computing, dynamic safe path planning, and multi-dimensional risk warning, reducing the false judgment rate and improving the safety and accuracy of emergency evacuation.
Smart Images

Figure CN122049997B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and public safety technology, and relates to a method and system for recognizing composite abnormal behaviors in fire protection and security coupled scenarios. Background Technology
[0002] Computer vision technology, as a key branch of artificial intelligence, has been widely applied in the field of public safety monitoring. Through the automatic analysis and understanding of video surveillance images, continuous monitoring of targets and events in a scene can be achieved. Especially in densely populated public places, image recognition technology plays a crucial role in effective fire safety management and daily security deployment, ensuring the safety of people's lives and property. This type of technology typically focuses on extracting valuable information from continuous video streams to assist or replace manual monitoring, achieving intelligent risk perception.
[0003] Currently, in intelligent video surveillance applications, fire safety solutions typically focus on detecting direct fire phenomena such as flames and smoke, or using methods like background subtraction to determine if static objects obstruct fire exits. In the realm of everyday security, solutions focus more on recognizing specific individual behaviors. For example, deep learning models can be used to identify key points in the human skeleton to determine if a fall has occurred, or tracking target trajectories can detect intrusions or loitering. These technologies mostly operate in their own independent scenarios, forming detection algorithms tailored to specific events.
[0004] However, existing technologies exhibit significant limitations when dealing with complex scenarios involving the coupling of fire protection and security: 1. Lack of a unified physical scale benchmark. Due to camera perspective distortion, the pixel variation at the same physical distance in a two-dimensional image exhibits non-linear differences. Therefore, directly extracting features in the pixel coordinate system makes it difficult to set universal thresholds using unified physical dimensions, thus limiting the consistency of spatial measurements and their applicability across different scenarios.
[0005] 2. Lack of dynamic topology reconstruction mechanism. Existing scene semantic models are mostly fixed configurations. Although they can detect static obstacles, they cannot map them to real-time changes in the scene's topology. When a passage is partially blocked, the system still uses the initial spatial guidance and cannot dynamically update the safe path vector, causing a misalignment between the judgment criteria for behaviors such as reverse travel and the actual physical constraints.
[0006] 3. Lack of multi-factor coupled computing mechanism. Existing anomaly identification is mostly an independent task, failing to integrate individual status, spatial attributes, and the movement trajectories of surrounding crowds. Due to the lack of multi-dimensional collaborative computing logic, the system struggles to achieve coordinated early warning of group risk evolution in complex situations such as individual falls combined with crowd convergence. Summary of the Invention
[0007] In order to overcome the above-mentioned defects of the prior art and to achieve the above objectives, the present invention proposes the following technical solution: a composite abnormal behavior recognition method in a fire protection and security coupled scenario, comprising: S1, performing coordinate transformation based on the calibration parameters of the monitoring camera, mapping the two-dimensional image coordinates to the top-view physical coordinate system, and fusing scene structure information to generate an initial scene geometric semantic model.
[0008] S2. During the period when no fire alarm signal is received, perform object detection and contour extraction to obtain the physical contour of the object left in the monitoring screen.
[0009] S3. Perform obstacle interference analysis, calculate the physical interference between the physical contour of the abandoned object and the moving area of the fire door in the scene geometric semantic model. When the physical interference exceeds the preset interference threshold, it is determined that a passage blockage event has occurred. Based on the physical position of the abandoned object, the scene geometric semantic model is dynamically updated to generate a dynamic semantic model.
[0010] S4. In response to the received fire alarm signal, perform evacuation behavior recognition based on the dynamic semantic model and extract the individual movement trajectory vectors of the evacuating crowd.
[0011] S5. Generate reverse behavior detection results by comparing the directional relationship between the individual motion trajectory vector and the safety guidance vector in the dynamic semantic model.
[0012] S6. Perform dense crowd risk monitoring in the evacuation bottleneck area identified by the dynamic semantic model, identify targets that meet the fall pattern criteria as potential fall targets, and analyze the displacement direction characteristics of the crowd around the potential fall targets.
[0013] S7. When the displacement direction characteristics meet the centripetal convergence condition, a stampede risk warning result is generated.
[0014] The second aspect of the present invention provides a composite abnormal behavior recognition system in a fire protection and security coupled scenario, comprising: a semantic mapping module, which performs coordinate transformation based on the calibration parameters of the monitoring camera, maps the two-dimensional image coordinates to the top-view physical coordinate system, and integrates scene structure information to generate an initial scene geometric semantic model.
[0015] The debris detection module performs debris detection and contour extraction during periods when no fire alarm signal is received, obtaining the physical contours of debris in the monitoring screen.
[0016] The dynamic reconstruction module performs obstacle interference analysis, calculates the physical interference between the physical contour of the abandoned object and the moving area of the fire door in the scene geometric semantic model, determines that a passage blockage event has occurred when the physical interference exceeds the preset interference threshold, and dynamically updates the scene geometric semantic model based on the physical position of the abandoned object to generate a dynamic semantic model.
[0017] The trajectory extraction module, in response to the received fire alarm signal, performs evacuation behavior recognition based on a dynamic semantic model and extracts the individual movement trajectory vectors of the evacuating crowd.
[0018] The reverse driving detection module generates reverse driving behavior detection results by comparing the directional relationship between the individual's motion trajectory vector and the safety guidance vector in the dynamic semantic model.
[0019] The fall detection module performs dense crowd risk monitoring in evacuation bottleneck areas identified by the dynamic semantic model, identifies targets that meet the fall pattern criteria as potential fall targets, and analyzes the displacement direction characteristics of people around potential fall targets.
[0020] The trampling warning module generates a trampling risk warning result when the displacement direction characteristics meet the centripetal convergence condition.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention maps video information to a unified top-down physical coordinate system and uses a unified physical dimension to quantify the channel blockage interference, the remaining passage width, and the individual motion trajectory vector. This method eliminates the image scale nonlinearity difference caused by camera perspective distortion, enabling the system to use a unified threshold standard for calculation under different fields of view, and improving the adaptability of the algorithm to cross-scene deployment.
[0022] (2) This invention maps structural information such as access control opening range and evacuation bottleneck area to physical space by constructing a dynamic geometric semantic model. When an obstacle appears in the physical passage, the system synchronously updates the passage topology and dynamically replans the safety guidance vector. The determination of reverse behavior is directly compared with the dynamically generated safety guidance vector, avoiding logical misjudgments caused by using fixed entrance and exit directions.
[0023] (3) The risk warning logic of this invention integrates individual status, spatial attributes, and the movement trajectory of surrounding crowds. For the warning of stampede risk, the system performs multi-dimensional coupling judgment based on the posture of the potential fall target, the evacuation bottleneck area where it is located, and the centripetal convergence characteristics of the surrounding crowd. Through the above multi-factor collaborative computing mechanism, the system outputs analysis results that conform to the mechanical logic of emergency evacuation, reducing the false alarm rate under single visual feature detection. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1This is a schematic diagram of the implementation steps of the method of the present invention.
[0026] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] Please see Figure 1 As shown, the present invention proposes a composite abnormal behavior recognition method in a fire protection and security coupled scenario, which includes: S1, performing coordinate transformation based on the calibration parameters of the monitoring camera, mapping the two-dimensional image coordinates to the top-view physical coordinate system, and fusing scene structure information to generate an initial scene geometric semantic model.
[0029] In a preferred embodiment, a coordinate transformation is performed based on the calibration parameters of the surveillance camera to map the two-dimensional image coordinates to a top-down physical coordinate system, and scene structure information is fused to generate an initial scene geometric semantic model, including:
[0030] S1.1 Obtain the internal parameters of the surveillance camera and its external parameters relative to the ground, and calculate the homography transformation matrix from the image plane to the ground physical plane;
[0031] S1.2 Identify the key structural points of the fire door in the image, and use the homography transformation matrix to transform the coordinates of the key structural points to the top-view physical coordinate system to obtain the physical position and physical dimensions of the door;
[0032] S1.3 In the top-view physical coordinate system, based on the physical position and physical size of the door, construct a physical sweep area that represents the opening range of the door leaf, and combine it with the preset safety exit position to generate the initial safety guide vector, and mark the initial evacuation bottleneck area according to the physical width of the passage, together forming the initial scene geometric semantic model.
[0033] Specifically, the goal of this step is to construct a top-down scene map containing critical security information, corresponding to real-world dimensions, using fixed surveillance camera footage—that is, the initial scene geometric semantic model. This model forms the basis for all subsequent analyses.
[0034] First, the calibration parameters of the surveillance camera are obtained. These parameters describe the camera's imaging characteristics, as well as its installation position and orientation in three-dimensional space. Calibration parameters are divided into two categories: internal parameters and external parameters. Internal parameters describe the camera's optical properties, such as the camera's focal length and the center point of the image sensor; external parameters describe the rotation and translation relationship between the camera's coordinate system and the world's physical coordinate system, such as the camera's height above the ground and its pitch angle. These parameters are typically calculated in one go using mature computer vision algorithms such as the Zhang Zhengyou calibration method, after taking multiple images and placing a calibration board of known size and pattern within the camera's field of view.
[0035] Using these internal and external parameters, a key mathematical tool can be derived: the homography transformation matrix. Since the ground in a surveillance scene is a plane, there exists a direct two-dimensional to two-dimensional projection mapping relationship from any point on the camera's image plane to its corresponding point on the physical ground plane. This relationship can be precisely described by a 3x3 homography transformation matrix. This transformation converts points in the pixel coordinate system of the image into points in the real-world top-down physical coordinate system, expressed in meters. The transformation relationship is expressed as follows:
[0036]
[0037] In the formula, Represents the coordinates of a two-dimensional image, where and These are the coordinates of the pixel in the horizontal and vertical directions of the image, respectively. This represents the corresponding coordinates of a pixel in the top-view physical coordinate system, where and The horizontal and vertical physical distances of pixels in the top-down physical coordinate system, in meters. The scale factor is used to normalize homogeneous coordinates. It is a 3x3 homography transformation matrix, the specific values of which are determined by the camera's intrinsic parameter matrix. And the rotation matrix in the external parameters and translation vector A joint decision was made. Based on 50 calibration experiments conducted in a laboratory environment on a camera installed at a height of 3 meters and a downward angle of 30 degrees, the matrix... The value is set to a robust mean result, ensuring the accuracy of the coordinate transformation.
[0038] After obtaining the homography transformation matrix, the next step is to process the scene structure information in the image, especially the fire door. Using image recognition algorithms, such as deep learning-based object detection models, the key structural points of the fire door are accurately located in the image, such as the four corners of the door frame or the two bottom endpoints of the door leaf. Assuming the coordinates of the two bottom endpoints of the fire door leaf in the image are... and Using the aforementioned homography transformation matrix Transform the two two-dimensional image coordinates into the top-view physical coordinate system to obtain the corresponding physical coordinates. and These two physical coordinate points precisely define the physical location of the door. The physical dimensions of the door, especially its width... This can be obtained by calculating the Euclidean distance between these two physical coordinate points:
[0039]
[0040] In the formula, This refers to the physical width of the fire door, measured in meters. and These are the coordinates of the two bottom endpoints of the door leaf in the top-down physical coordinate system.
[0041] Finally, in the top-down physical coordinate system, based on the obtained physical information of the door, an initial scene geometric semantic model is constructed. This model contains three core elements: the first is the physical sweep area, which characterizes the spatial range swept by the door when fully open. For a single door, if its hinge position is in... The door width is Its physical sweep area is a region that is With the center of the circle, A sector-shaped area with a radius of 0.5 and an opening angle of 90 degrees. This area must be kept clear in an emergency.
[0042] The second element is the initial safety guidance vector. This is combined with the pre-set safety exit location coordinates in the system. Vectors pointing to safety exits can be generated at different locations within the passage. For example, at any point... Its initial safety guidance vector It can be calculated as follows:
[0043]
[0044] This vector provides a benchmark for subsequent determination of the direction of personnel evacuation.
[0045] The third element is the initial evacuation bottleneck area. Based on building fire safety codes and the physical width of the on-site passageways, a width threshold for one person's passageway capacity is set. For example, based on analysis of passageway pedestrian flow data measured by 200 sets of industrial sensors, the minimum comfortable width for two people to walk side-by-side is 1.2 meters. When the measured physical width of the passageway is less than this threshold, that passageway section is marked as the initial evacuation bottleneck area.
[0046] By integrating the three pieces of information—the physical sweep area, the initial safety guidance vector, and the initial evacuation bottleneck area—a rich initial scene geometric semantic model is formed, which provides a static scene benchmark for subsequent dynamic analysis.
[0047] For example, suppose a surveillance camera is installed on the corridor ceiling, and its image resolution is 1920x1080 pixels. After calibration, the resulting homography transformation matrix... for:
[0048]
[0049] In the surveillance footage, image analysis identified two key structural points at the bottom of the fire door leaf, with their two-dimensional image coordinates as follows: and .
[0050] To obtain the physical location and dimensions of the door, these two coordinate points are substituted into the transformation formula.
[0051] For point :
[0052]
[0053] After normalization, that is, each term is divided by the scale factor. To obtain the coordinates in the top-view physical coordinate system for rice.
[0054] For point :
[0055]
[0056] After normalization, the coordinates in the top-view physical coordinate system are obtained. for rice.
[0057] The physical dimensions (width) of the door. The calculation is as follows:
[0058]
[0059] The physical location of the door is determined by a point. and Sure.
[0060] Next, an initial scene geometric semantic model is constructed in the top-down physical coordinate system.
[0061] Assuming the hinges of the fire door are in Then its physical sweep area is a region consisting of... A 90-degree sector with a center of 0.87 meters and a radius of 0.87 meters.
[0062] Pre-set safety exit locations are For a point in the corridor Its initial safety steering vector is .
[0063] Measurements showed that the physical width of the passageway containing the fire door was 1.4 meters. The set width threshold was 1.2 meters; because 1.4 meters is greater than 1.2 meters, this area was not marked as an initial evacuation bottleneck.
[0064] Ultimately, the initial scene geometric semantic model is formed by the fan-shaped physical sweep area, a series of pre-computed initial safety guidance vectors, and the initial evacuation bottleneck area marking information within the scene.
[0065] S2. During the period when no fire alarm signal is received, perform object detection and contour extraction to obtain the physical contour of the object left in the monitoring screen.
[0066] Specifically, this step is performed under normal system operation. Its core task is to identify abandoned objects in the monitored scene that have changed from dynamic to static, and to obtain the actual outline of these abandoned objects on the ground plane. This process is defined as abandoned object detection and outline extraction under normal conditions, ultimately producing the physical outline of the abandoned objects.
[0067] First, the system distinguishes between background and foreground in the scene. It employs background modeling algorithms (such as Gaussian mixture models or multi-frame averaging) to process continuous video images and establish a dynamic background model. When a new frame of video image is acquired, it is compared pixel-by-pixel with the background model. If the difference between a pixel's feature vector (such as grayscale value or color feature) and the corresponding background pixel is greater than a preset background difference threshold, the pixel is marked as foreground; otherwise, it is marked as background. The set of pixels marked as foreground constitutes a foreground mask map, used to represent active or newly appearing targets in the scene.
[0068] However, foreground objects are not the same as abandoned objects; for example, a pedestrian walking normally is also in the foreground. Therefore, further analysis is needed to identify objects that change from dynamic to static states. The system continuously tracks each connected region in the foreground mask image, i.e., each individual target object, using a target tracking algorithm (such as a centroid-matching-based tracking algorithm). For each tracked target, the system calculates its center point in the image. If the center point of a target remains stable for a continuous period of time, i.e., the displacement change is less than a preset motion threshold, the system starts a timer. If the target remains stationary for a longer period than the abandoned object determination time, it is identified as an abandoned object. This determination of a stationary state can be expressed by the following formula:
[0069]
[0070] In the formula, and These are the centroids of the target in the image at the current moment. and the previous moment The pixel coordinates. The motion threshold, in one specific embodiment, can be set to 2 pixels. This value, based on 1000 hours of surveillance video analysis, can effectively filter out false motion caused by camera sensor noise or slight environmental disturbances. The time limit for determining abandoned objects can be set according to security management regulations, for example, 300 seconds.
[0071] Once the target is identified as a remnant, the system uses a contour extraction algorithm to obtain the boundary of the region corresponding to it in the foreground mask image, forming a sequence of pixel coordinates. The closed image outline formed This represents the total number of boundary pixels that constitute the image contour. The contour represents the shape of the remnant in the two-dimensional image.
[0072] Finally, the homography transformation matrix calculated in step S1 is used. The system will close each pixel on the image outline. Transform each point to the top-down physical coordinate system to obtain the corresponding physical coordinate points. The transformation process is as follows:
[0073]
[0074] In the formula, It is the first one on the image outline Let i = 1, 2, ..., n be the homogeneous pixel coordinates of n points. It is the first one on the image outline The homogeneous physical coordinates corresponding to each point It is the first one on the image outline Scale factor for each point. All transformed physical coordinate points. The resulting closed polygon is the final physical outline of the remains. This physical outline represents the actual area and shape of the remains on the ground, measured in meters.
[0075] It should be noted that, It is a homography transformation matrix Homogeneous coordinates of pixels After performing matrix multiplication, the third element of the resulting intermediate column vector is obtained. The purpose of this calculation is to convert the projected 3D homogeneous coordinates into 2D planar coordinates, and its value dynamically changes depending on the pixel coordinates of the input image. For example, when substituting a specific matrix... When performing a transformation operation on a pixel, if the initially calculated intermediate column vector is... If so, the third term of the vector is extracted, and the value of the scale factor corresponding to the pixel is determined to be 1.046.
[0076] For example, continuing the scenario from step S1, the surveillance camera and homography transformation matrix used are... Remain unchanged:
[0077]
[0078] During the obstacle interference analysis under normal conditions, the monitoring system detected a sanitation worker pushing a cleaning cart into the frame, stopping in the area in front of the fire door, then unloading a square cardboard box from the cart and placing it on the ground before pushing the cart away.
[0079] The system first models the background, identifying the cleaner, cleaning cart, and cardboard box as foreground elements. Once the cleaner and cleaning cart leave, only the stationary cardboard box remains in the foreground. The system tracks the centroid pixel coordinates of the cardboard box and finds that its centroid stabilizes within a stable position over a continuous 10 seconds. If the displacement change is less than 2 pixels, the static condition is met. The system starts timing, and after the cardboard box remains stationary for 300 seconds, the system officially classifies it as abandoned property.
[0080] At this point, the system extracts the outline of the cardboard box in the image, obtaining the pixel coordinates of its four bottom vertices as follows: , , and .
[0081] To obtain the physical contours of the relic, the system utilizes a homography transformation matrix. Perform coordinate transformations on these four vertices. For each vertex... Perform matrix multiplication:
[0082] Extract the third term from the middle column vector as the scaling factor. After normalization, the physical coordinates are obtained. rice, Meters. That is, physical coordinates. for rice.
[0083] Similarly, the remaining vertices are transformed and normalized for calculation:
[0084] For vertices The physical coordinates are calculated. rice.
[0085] For vertices The physical coordinates are calculated. rice.
[0086] For vertices The physical coordinates are calculated. rice.
[0087] These four physical coordinate points , , , Connecting them, they form a quadrilateral. This quadrilateral is the physical outline of the remnant (cardboard box) in the top-down physical coordinate system. It accurately represents the actual position and range occupied by the cardboard box on the ground and can be used for subsequent interference calculations.
[0088] S3. Perform obstacle interference analysis, calculate the physical interference between the physical contour of the abandoned object and the moving area of the fire door in the scene geometric semantic model. When the physical interference exceeds the preset interference threshold, it is determined that a passage blockage event has occurred. Based on the physical position of the abandoned object, the scene geometric semantic model is dynamically updated to generate a dynamic semantic model.
[0089] In a preferred embodiment, the scene geometric semantic model is dynamically updated based on the physical location of the remnants to generate a dynamic semantic model, including:
[0090] S3.1 When a passage blockage event is determined to occur, the physical outline of the remaining object is marked as a dynamic obstacle area in the top-view physical coordinate system;
[0091] S3.2 In the top-view physical coordinate system, path replanning is performed with the dynamic obstacle area as a constraint. The avoidance path from the upstream of the dynamic obstacle area to the safety exit is calculated, and the direction of the avoidance path at the key position is extracted and normalized to generate a dimensionless unit vector as the updated safety guidance vector.
[0092] S3.3 Calculate the remaining passage width between the edge of the dynamic obstacle area and the wall boundaries on both sides of the passage. Mark the sections with remaining passage width less than the preset human passage width threshold as the updated evacuation bottleneck area. The dynamic obstacle area, the updated safety guidance vector, and the updated evacuation bottleneck area constitute the dynamic semantic model.
[0093] In a further preferred embodiment, in the top-view physical coordinate system, path replanning is performed with dynamic obstacle areas as constraints to calculate an avoidance path from the upstream of the dynamic obstacle area to the safety exit. Specifically, a graph search algorithm is used to discretize the top-view physical coordinate system into a grid, and the grid occupied by the dynamic obstacle area is set as an impassable node. The target passable path from the preset starting grid to the safety exit grid is searched as the avoidance path.
[0094] Specifically, after obtaining the initial scene geometric semantic model established in step S1 and the physical outline of the debris identified in step S2, the core of this step is to determine whether the debris constitutes an actual blockage of the fire lane, and after confirming the blockage, update the original scene model in real time to reflect the latest safety status on site.
[0095] First, the degree to which the debris obstructs the normal opening of the fire door is quantified. Using computational geometry algorithms (such as the Sutherland-Hodgman polygon clipping algorithm), the overlap area between the physical contour of the debris (a polygonal region) and the physical sweep area of the fire door (a sector-shaped region) defined in the initial scene geometric semantic model is calculated in the top-view physical coordinate system. This overlap area is defined as the physical interference. Its calculation can be expressed as:
[0096]
[0097] In the formula, This represents the amount of physical interference, measured in square meters. The polygonal region defined by the physical outline of the remnant. This represents the physical sweep area where the fire door is open. It is a function for calculating the area of a geometric region. This represents the intersection operation of two regions.
[0098] The calculated physical interference needs to be compared with a preset interference threshold. The interference threshold is a minimum tolerable obstruction area set according to fire safety regulations to avoid false alarms caused by negligible obstacles or calculation errors. Based on the analysis of fire drill data in public places, the interference threshold is set at 0.01 square meters. When the calculated physical interference exceeds this threshold, the system determines that a passageway blockage event has occurred.
[0099] Once a channel congestion event is detected, the system will immediately initiate a dynamic update process for the scenario model, generating a completely new dynamic semantic model. This update process involves three coordinated operations.
[0100] First, the physical outlines of the debris that caused this passage blockage are marked in a top-down physical coordinate system, forming a dynamic obstacle area. This area is treated as a fixed, impassable obstacle in the new model.
[0101] Second, due to the emergence of dynamic obstacle zones, the original initial safety guidance vector may no longer be effective, as it may point to obstacles. Therefore, path replanning is required. The system employs a graph search algorithm, such as the A* algorithm, to accomplish this task. Specifically, the entire top-down physical coordinate system is discretized into a fine-grained grid map, with each grid cell representing a small square region. All grid cells covered by the dynamic obstacle zone are set as impassable nodes. Subsequently, the algorithm searches for one or more target paths that bypass the impassable nodes, starting from one or more preset starting grids upstream of the dynamic obstacle zone and ending at the grid containing the safety exit. This new path is the avoidance path. The updated safety guidance vector can be extracted from the avoidance path. For any point in the scene, its updated safety guidance vector is the direction vector pointing from that point to the next path point on the nearest point on the avoidance path, and it is normalized to become a dimensionless unit vector.
[0102] Third, the system reassesses the passage's capacity. The system retrieves the physical coordinates of the fixed boundaries (such as wall outlines) of the passage pre-stored in the initial scene geometric semantic model, calculates the minimum Euclidean distance between the edge vertices of the dynamic obstacle area and these fixed boundaries, and obtains the remaining passage width. This width is then compared to a preset human passage width threshold. The human passage width threshold is the minimum width required to ensure a single person can pass quickly; based on ergonomic data, this threshold is set to 0.8 meters. If the remaining passage width of a certain section is less than 0.8 meters, then that area is marked as an updated evacuation bottleneck area.
[0103] Ultimately, the newly marked dynamic obstacle areas, the updated safety guidance vectors obtained through path replanning, and the reassessed updated evacuation bottleneck areas together constitute the dynamic semantic model, providing accurate and real-time scenario information guidance for subsequent fire emergency evacuation.
[0104] For example, the scenario and data from the preceding steps are used. In step S1, the physical sweep area of the fire door has been determined as a point... A sector-shaped area with a center and a radius of 0.88 meters. In step S2, the physical outline of a remnant is identified, with its vertex being... , , and .
[0105] The system calculates the overlap area between the physical contour of the remaining object and the physical sweep area of the fire door, and obtains the physical interference. The area is 0.12 square meters. The set interference threshold is 0.01 square meters. Because... The system determines that a channel blockage event has occurred.
[0106] The system then begins to generate a dynamic semantic model.
[0107] First, from the vertex , , , The resulting quadrilateral region is marked as the dynamic obstacle region.
[0108] Next, path replanning is performed. The top-down physical coordinate system is divided into a 0.1m x 0.1m grid. Grids covering dynamic obstacle areas are set as impassable. The path is then replanned using a preset starting grid upstream of the obstacle. Starting with the safety exit grid With the destination as the endpoint, a graph search algorithm is executed. The avoidance path found by the algorithm is no longer a straight line, but a curve that bypasses the dynamic obstacle region. For example, the path may pass through point... and To avoid obstacles. For points that were originally near an obstacle and whose initial safety guidance vector pointed to the right, their updated safety guidance vector will now point upwards, guiding people to detour.
[0109] Finally, calculate the remaining passage width. Assume the other wall of the passage is a straight line in the top-view physical coordinate system. Meters. The vertex of the dynamic obstacle zone closest to the wall is... The distance from this point to the wall is... Meters. This distance is the remaining passage width. The set threshold for human passage width is 0.8 meters. Because... Therefore, the narrow passage between the dynamic obstacle area and the wall is marked as the updated evacuation bottleneck area.
[0110] Ultimately, this newly labeled dynamic obstacle region, the updated set of safety guidance vectors generated around it, and the newly identified updated evacuation bottleneck region together constitute the dynamic semantic model.
[0111] S4. In response to the received fire alarm signal, perform evacuation behavior recognition based on the dynamic semantic model and extract the individual movement trajectory vectors of the evacuating crowd.
[0112] In a preferred embodiment, extracting the individual movement trajectory vectors of the evacuated crowd includes:
[0113] S4.1 Perform target detection and tracking on the surveillance video sequence containing evacuated crowds to obtain the bounding rectangles of multiple human targets in consecutive image frames;
[0114] S4.2 For each tracked human target, select the bottom feature points of its bounding rectangle, and use the homography transformation matrix to map the bottom feature points to the top-view physical coordinate system to obtain the sequence of physical location points of the human target.
[0115] S4.3. Based on the changes in the sequence of physical location points over consecutive timestamps, calculate the individual motion trajectory vector corresponding to each human target.
[0116] Specifically, when the system receives an externally triggered fire alarm signal, it will immediately switch from normal monitoring to emergency evacuation mode. In this mode, the core objective of this step is to perform real-time analysis of the evacuating crowd in the monitoring screen based on the previously established dynamic semantic model (or the initial scene geometric semantic model under unobstructed conditions), and to calculate an individual motion trajectory vector representing the direction and speed of movement for each person.
[0117] First, the system uses a target detection model (such as YOLO or Faster R-CNN based on convolutional neural networks) to process each frame of the video, obtaining the bounding boxes of each human target in the monitored image. Then, a multi-target tracking algorithm (such as DeepSORT or SORT) is invoked to perform data association based on the intersection-over-union (IoU) and appearance features of the human targets across consecutive image frames, assigning identity identifiers to each human target, thereby extracting the sequence of bounding boxes of each human target corresponding to its identity in consecutive image frames.
[0118] After obtaining the bounding box sequence, the image coordinates need to be transformed to real-world coordinates. Since the human body is three-dimensional, while homography transformation is planar, a feature point that accurately represents the person's position on the ground needs to be selected. The midpoint of the bottom edge of the bounding box is an ideal choice because it typically corresponds to the position of the person's feet on the ground plane. For each tracked human target, the system extracts the midpoint of the bottom edge of the bounding box in each frame as the feature point representing the contact between the human body and the ground. Let the top-left corner pixel coordinates of the bounding box be... The bottom right pixel coordinates are Then the pixel coordinates of the bottom feature point Calculated as Subsequently, the homography transformation matrix calculated in step S1 is used. By mapping this two-dimensional image coordinates to a top-down physical coordinate system, the physical location of the human target at that moment can be obtained. Over time, the continuous physical location points constitute a sequence of physical location points for the human target.
[0119] Finally, the individual motion trajectory vector describing its motion state is calculated based on the sequence of physical location points. This vector is essentially a velocity vector in the physical world, containing both speed and direction of motion. It can be obtained by taking the displacement change of the human target at two consecutive points in the sequence of physical location points within a short time window and dividing it by the corresponding time interval. The calculation formula is as follows:
[0120]
[0121] In the formula, It is the calculated vector of the individual's motion trajectory, and its unit is meters per second. The human target at the current moment The physical location coordinates. The human target at the previous moment The physical location coordinates. It is the time difference between two consecutive timestamps; for example, for a video with a frame rate of 10 frames per second, its value is 0.1 seconds. The vector calculated by this formula accurately reflects the instantaneous movement of evacuees on the ground.
[0122] For example, suppose a fire alarm signal is triggered, and the system enters the evacuation phase behavior recognition phase. The homography transformation matrix used... Maintain consistency with the preceding steps:
[0123]
[0124] The video processing frame rate is 10 frames per second, corresponding to a time interval. Second.
[0125] The system performs target detection and tracking on the surveillance video sequence, assigning an identification identifier "ID-15" to a human target. In the... The bounding rectangle coordinates of frame ID-15 are the top left corner. bottom right corner In the first The frame, whose outer rectangle is updated to the top left corner. bottom right corner .
[0126] Next, the system selects the bottom feature points of the bounding rectangle for ID-15 and performs coordinate transformation.
[0127] In the The pixel coordinates of the bottom feature point in the frame are: Using the homography transformation matrix Perform matrix multiplication:
[0128]
[0129] Extracting scale factors Perform normalization to obtain the mapped physical location points. rice.
[0130] In the The pixel coordinates of the bottom feature point in the frame are: Similarly, performing mapping and normalization calculations yields the physical location point. rice.
[0131] Finally, based on these two consecutive physical location points, the individual motion trajectory vector of ID-15 at that moment is calculated. :
[0132]
[0133] Calculation results show that the individual motion trajectory vector of human target ID-15 is The speed is meters per second, meaning he is moving at a speed of 0.4 meters per second horizontally and 0.12 meters per second vertically. This vector data will be used for subsequent behavioral analysis.
[0134] S5. Generate reverse behavior detection results by comparing the directional relationship between the individual motion trajectory vector and the safety guidance vector in the dynamic semantic model.
[0135] In a preferred embodiment, the reverse behavior detection result is generated by comparing the directional relationship between the individual motion trajectory vector and the safety guidance vector in the dynamic semantic model, including:
[0136] S5.1 Obtain the updated safety guidance vector associated with the current individual's physical location in the dynamic semantic model;
[0137] S5.2 Calculate the dot product of the individual motion trajectory vector and the updated safety guidance vector;
[0138] S5.3 If the dot product value remains negative within a preset time duration threshold and the individual's motion trajectory vector crosses the predefined security alert boundary, then it is determined that the individual has committed reverse intrusion behavior, and a reverse behavior detection result is generated.
[0139] Specifically, this step aims to use the real-time movement information of each evacuee obtained in step S4 to compare with the evacuation guidance information in the dynamic semantic model established in step S3, thereby automatically identifying reverse behavior during emergency evacuation.
[0140] For each tracked human target, the system first finds the evacuation guidance information that best matches its current location from the dynamic semantic model. The system then obtains the physical location of the human target at the current moment. Then, in the updated safety guidance vector field defined by the dynamic semantic model, the updated safety guidance vector associated with the location point is obtained through a retrieval algorithm (such as bilinear interpolation). This vector represents the appropriate evacuation direction that the system plans based on the site conditions (such as obstacles) at that location.
[0141] After obtaining the two key vectors—the actual direction of movement and the intended direction of movement—the directional relationship is quantified by calculating their dot product. The dot product is an effective mathematical tool for measuring the similarity between the directions of two vectors. The calculation method is as follows:
[0142]
[0143] In the formula, It is the individual motion trajectory vector calculated from step S4. It is an updated security guidance vector obtained from the dynamic semantic model. It is the calculated dot product value, which is a scalar. If A positive value indicates that the angle between the individual's movement direction and the planned evacuation direction is less than 90 degrees; if A negative value indicates that the angle between the individual's movement direction and the planned evacuation direction is greater than 90 degrees; if A value of 0 indicates that the angle between the individual's movement direction and the planned evacuation direction is 90 degrees.
[0144] However, a single instantaneous negative dot product value is insufficient to determine reverse movement, as individuals may only briefly adjust their posture or avoid the obstacle. To ensure detection accuracy, the system introduces two additional constraints: First, temporal persistence: the system monitors whether the dot product value of a target individual remains continuously negative. This duration is set to 2 seconds based on the analysis of behavioral patterns in 500 simulated evacuation videos. This means that only when a person's direction of movement remains opposite to the safety guidance for more than 2 consecutive seconds is it considered potential reverse movement. Second, spatial boundary: the system predefines several security warning boundaries in a top-down physical coordinate system. These boundaries are typically set at critical nodes along the evacuation route, such as the entrance from the safety passage back to the danger zone.
[0145] Finally, when the dot product of the individual's trajectory vector and the updated security guidance vector remains negative, and the trajectory crosses the predefined security perimeter during this period of negativity, the system determines that the individual has committed intrusion behavior and generates an intrusion detection result. This result will include the individual's identification, the specific time and physical location of the intrusion, enabling security personnel to respond quickly.
[0146] For example, continuing with the scenario from the previous steps, the system is tracking a human target "ID-15". At a certain moment, the physical location of ID-15 is... rice.
[0147] The system queries the dynamic semantic model to obtain the updated safety guidance vector associated with that location point. This vector represents the planned evacuation direction.
[0148] In the subsequent 0.1-second interval, the physical location of ID-15 shifted to... Meters. Based on this, the vector of its individual movement trajectory is calculated as:
[0149]
[0150] This vector indicates that ID-15 is moving in the opposite direction to what it was before.
[0151] The system then calculates the dot product of the two vectors. :
[0152]
[0153] The calculated dot product value was negative. The system continued to monitor, and for the next 2 seconds (corresponding to 20 video frames), the calculated dot product value remained negative.
[0154] Meanwhile, the system pre-stored in The coordinates of the security perimeter boundary at the specified distance. During the ID-15's reverse movement, its physical X-coordinate changed from 4.07 meters to 3.99 meters, and its trajectory crossed the security perimeter boundary.
[0155] Because the dot product value remained negative for 2 seconds and the individual's trajectory vector crossed the predefined security perimeter, the system ultimately determined that ID-15 had committed a reverse intrusion, generated a reverse intrusion detection result, and recorded ID-15's coordinates. An incident of going against the flow of traffic across the border.
[0156] S6. Perform dense crowd risk monitoring in the evacuation bottleneck area identified by the dynamic semantic model, identify targets that meet the fall pattern criteria as potential fall targets, and analyze the displacement direction characteristics of the crowd around the potential fall targets.
[0157] In a preferred embodiment, identifying targets that meet the fall pattern criteria as potential fall targets includes:
[0158] S6.1 Within the evacuation bottleneck area, continuously monitor the shape and height parameters of the bounding rectangle of each human target.
[0159] S6.2 When the morphological proportion parameters of a human target change from the set upright characteristic range to the fallen characteristic range within a preset time window, and the decrease in its position height parameter exceeds the preset height threshold, it is determined that the human target has fallen.
[0160] S6.3 Mark the physical location of the target that has fallen as the potential center of mass of the fall, and identify the human target as a potential fall target.
[0161] In a further preferred embodiment, analyzing the displacement direction characteristics of people surrounding the potential fall target includes:
[0162] S6.4. Delineate a circular monitoring area in the top-view physical coordinate system with the potential center of gravity of the fall as the center;
[0163] S6.5. Obtain other human targets that enter the circular monitoring area within the set time window, and calculate the individual motion trajectory vectors of these human targets;
[0164] S6.6 Analyze the directionality of each individual's motion trajectory vector relative to the potential center of mass of the fall to obtain the displacement direction characteristics of the surrounding crowd.
[0165] Specifically, this step focuses on one of the high-risk aspects of fire evacuation: falls that may occur in crowded areas and the resulting chain reaction of stampedes. The system focuses its monitoring on evacuation bottleneck areas already identified in the dynamic semantic model, performing risk monitoring of dense crowds. This monitoring consists of two main parts: first, identifying people who may have fallen, and then analyzing the dynamic reactions of the people around them.
[0166] The first part identifies targets that meet the fall pattern criteria as potential fall targets. This process is achieved by continuously monitoring the bounding box of each human target located within the evacuation bottleneck area. The system focuses on changes in two key parameters: the morphological proportion parameter and the positional height parameter. Morphological proportion parameter Defined as the height of the circumscribed rectangle With width The ratio of .
[0167]
[0168] In the formula, and All units are pixels. For an upright human target, this ratio falls within a set upright characteristic range, which is set, for example, to 2.0 to 4.0 based on sample statistical analysis. When a fall occurs, the height of the outer rectangle decreases and the width increases, and the shape proportion parameter decreases to a set falling characteristic range, which is set, for example, to 0.3 to 1.0.
[0169] Simultaneously, the system extracts the position height parameter, which is set as the vertical pixel coordinate of the centroid of the circumscribed rectangle. When a person falls, their center of gravity drops rapidly, which is reflected in the image as a sudden drop in the positional height parameter. When the system detects that the morphological proportion parameter of a human target changes from an upright feature value to a fallen feature value within a preset time window (e.g., within 0.5 seconds), and the drop in its positional height parameter exceeds a preset height threshold, the system determines that the human target has experienced a postural fall.
[0170] Once a fall is detected, the system immediately identifies the person as a potential fall target. Simultaneously, the physical location of the fall (calculated using the method in step S4) is marked as a special point in the top-view physical coordinate system, called the potential fall centroid. .
[0171] The second part analyzes the displacement direction characteristics of people around the potential fall target. This process begins immediately after the potential fall target is identified. The potential fall centroid is then used as the marker. Centered on the target, the system defines a virtual circular monitoring area in the top-down physical coordinate system. The radius of this area is set based on the range of mutual influence between individuals in crowd dynamics research, specifically 2.0 meters, to ensure coverage of surrounding people who may interact with the person who has fallen.
[0172] Subsequently, the system will continuously monitor all other human targets entering this circular monitoring area within a set time window (e.g., the next 5 seconds). For each human target entering the area, the system will calculate its real-time individual motion trajectory vector, just as in step S4. .
[0173] Finally, the system analyzes the directionality of the individual motion trajectory vector of each surrounding target relative to the potential fall centroid. Specifically, for any surrounding target, the system calculates a trajectory vector from its current position... Pointing to the potential center of mass of the fall vector Then, through comparison and The direction of movement is used to determine whether people in the vicinity are approaching, moving away from, or going around the fallen person. The set of results from this directional analysis together constitutes the displacement direction characteristics of the surrounding crowd.
[0174] For example, within the updated evacuation bottleneck area marked in step S3, the system is conducting risk monitoring of densely populated areas. The physical coordinates of this area range from [missing information]. Coordinates at Rice to Between meters Coordinates at Rice to Between meters.
[0175] The system is tracking a human target identified as "ID-28" within the area. At time... The bounding rectangle of ID-28 has a height of 120 pixels and a width of 35 pixels; its proportions are as follows. It falls within the defined upright feature range [2.0, 4.0]. Its position height parameter is the centroid Y-coordinate. Pixel.
[0176] At any moment Within a preset 0.5-second time window, the bounding rectangle of ID-28 suddenly changes to a height of 40 pixels and a width of 90 pixels, adjusting the aspect ratio parameters. It falls within the set knockdown feature range [0.3, 1.0]. Simultaneously, its position height parameter changes to... Pixel (in the image coordinate system, the larger the Y value, the lower the position). Calculate its drop magnitude as... Pixels. The set height drop threshold is 50 pixels. Due to the variation of shape proportion parameters across ranges and the drop magnitude... The system determined that ID-28 had experienced a postural fall.
[0177] The system identifies ID-28 as a potential fall target and places it in... Physical location at time The meter is marked as the potential center of mass for a fall.
[0178] Next, with Centered on the target, the system delineates a circular monitoring area with a radius of 2.0 meters in the physical coordinate system. Over the next 5 seconds, the system detects three human targets, "ID-31," "ID-32," and "ID-33," entering this area.
[0179] For ID-31, its current physical location is Its individual motion trajectory vector is The vector pointing from its position to the potential center of mass of the fall is... The system calculates the cosine of the angle between two vectors: The set proximity judgment cosine threshold is 0.707. Because... The system determined that ID-31's movement trend was toward the person who had fallen.
[0180] Similarly, for ID-32 and ID-33, the system calculates the cosine of their corresponding angles and determines that the movement trend of ID-32 is to detour, while the movement trend of ID-33 is to approach.
[0181] These analysis results, "ID-31 approaching," "ID-32 detouring," and "ID-33 approaching," collectively constitute the displacement direction characteristics of the surrounding crowd at this moment, providing quantitative input for the next step of stampede risk warning.
[0182] S7. When the displacement direction characteristics meet the centripetal convergence condition, a stampede risk warning result is generated.
[0183] In a preferred embodiment, a stampede risk warning result is generated when the displacement direction characteristics satisfy the centripetal convergence condition, including:
[0184] S7.1 Count the number of human targets around the potential center of mass of a fall whose individual motion trajectory vector continuously points to the center of mass within a continuous preset time frame.
[0185] S7.2 If the number of human targets in the surrounding area exceeds the set threshold for the number of people to gather, then the displacement direction feature is determined to meet the centripetal gathering condition.
[0186] S7.3 When there are potential falling targets and the centripetal convergence condition is met, a stampede risk warning result is generated.
[0187] Specifically, this step is the final decision-making stage for risk monitoring in densely populated areas. After potential fall targets have been identified and the displacement direction characteristics of the surrounding crowd have been analyzed in step S6, this step uses quantitative analysis of these displacement direction characteristics to ultimately determine whether there is an imminent risk of stampede, and generates an early warning when specific conditions are met.
[0188] The core of this decision-making process is determining whether the movement of the surrounding crowd constitutes a centripetal convergence condition. To achieve this, the system first statistically analyzes the displacement direction features obtained in step S6. Specifically, the system counts the number of surrounding human targets whose individual motion trajectory vectors continuously point towards the potential center of mass within a continuous preset time window. The movement of a surrounding human target is determined to be "pointing" towards the potential center of mass when its individual motion trajectory vector... The vector pointing from the individual's current position to the potential center of mass of the fall. The angle between the two vectors is less than a preset directional angle threshold. This judgment can be achieved by calculating the unit vector dot product of the two vectors:
[0189]
[0190] In the formula, It is the cosine of direction consistency, obtained by dividing the dot product of two vectors by the product of their magnitudes, and its value is in the dimensionless real number range of [−1,1]. The mathematical symbol used to calculate the magnitude of a vector. When... If the value is greater than or equal to a preset cosine threshold, the individual's movement is considered directional. The system continuously checks this condition within a preset time window (set to 1.5 seconds based on crowd reaction speed). Only individuals who continuously meet the condition are counted, and the total number of individuals is recorded as follows. .
[0191] Next, the system will count the number of human targets in the surrounding area. Compared with the preset clustering threshold A comparison is needed. The gathering number threshold is a key critical value for determining whether crowd behavior has shifted from individual reactions to dangerous group aggregation. Based on the analysis of simulated crowd behavior data from 100 public safety incidents, when the number of people gathering towards a fallen person reaches 3 or more, the risk of stampede increases dramatically. Therefore, the gathering number threshold... It is set to 2. If Exceeded The system then determines that the current displacement direction characteristics meet the centripetal convergence condition.
[0192] Finally, the system performs a final risk assessment. Generating a stampede risk warning requires two preconditions to be met simultaneously: first, a potential fall target has been identified in step S6; second, the centripetal convergence condition determined in this step is satisfied. When both conditions are met, the system confirms a high risk of stampede and immediately generates a stampede risk warning result. This result includes information such as the warning level, time of occurrence, precise physical coordinates of the potential fall centroid, and the number of people gathered, providing decision support for on-site security and emergency command.
[0193] For example, continuing from step S6. There exists a system located at physical coordinates... The potential center of gravity of the fall (ID-28) and three human targets ID-31, ID-32 and ID-33 within a circular monitoring area with a radius of 2.0 meters.
[0194] The system begins performing statistical analysis. Within the next 1.5-second preset time window, the system continuously analyzes the directional movement of the three human targets.
[0195] For ID-31, its individual motion trajectory vector and the vector pointing to the potential center of mass of the fall Calculated direction consistency cosine value Within this time window, the mean value is 0.95, which is greater than the preset cosine threshold of 0.707. Therefore, ID-31 meets the directional motion condition and is counted.
[0196] For ID-32, its individual motion trajectory vector indicates that it is orbiting, and the calculated... The mean is 0.12, which is less than 0.707. Therefore, ID-32 does not meet the directional motion condition and is not counted.
[0197] For ID-33, its individual motion trajectory vector and the vector pointing to the potential center of mass of the fall Calculated direction cosine value The mean value within this time window is 0.89, which is greater than 0.707. Therefore, ID-33 meets the directional motion condition and is counted.
[0198] In addition, a new target ID-34 entered the monitoring area. Its directional consistency cosine value was calculated to be greater than 0.707, which met the directional motion condition and was counted.
[0199] The final statistics showed that the number of human targets around the potential center of gravity of the fall was continuously pointed towards the individual's motion trajectory vector. .
[0200] The system then makes comparisons based on the set threshold for the number of clusters. Due to statistical quantity Exceeding the clustering threshold The system determines that the displacement direction characteristics satisfy the centripetal convergence condition.
[0201] Finally, the system makes a decision. A potential falling target (ID-28) exists in the current scenario, and the centripetal convergence condition is met. Both necessary conditions have been met. Therefore, the system immediately generates a stampede risk warning result, which includes: the physical coordinates of the potential falling centroid. The quantitative data indicators include 1 person who fell and 3 people who gathered together.
[0202] Please see Figure 2 As shown, the second aspect of the present invention provides a composite abnormal behavior recognition system in a fire protection and security coupled scenario, comprising: a semantic mapping module, a legacy detection module, a dynamic reconstruction module, a trajectory extraction module, a reverse movement detection module, a fall recognition module, and a stampede warning module.
[0203] The semantic mapping module performs coordinate transformation based on the calibration parameters of the surveillance camera, maps the two-dimensional image coordinates to the top-view physical coordinate system, and integrates scene structure information to generate an initial scene geometric semantic model.
[0204] The abandoned object detection module performs abandoned object detection and contour extraction during periods when no fire alarm signal is received, and obtains the physical contour of abandoned objects in the monitoring screen.
[0205] The dynamic reconstruction module performs obstacle interference analysis, calculates the physical interference between the physical contour of the abandoned object and the moving area of the fire door in the scene geometric semantic model, determines that a passage blockage event has occurred when the physical interference exceeds a preset interference threshold, and dynamically updates the scene geometric semantic model based on the physical position of the abandoned object to generate a dynamic semantic model.
[0206] The trajectory extraction module, in response to the received fire alarm signal, performs evacuation behavior recognition based on a dynamic semantic model to extract individual movement trajectory vectors of the evacuating crowd.
[0207] The reverse movement detection module generates reverse movement detection results by comparing the directional relationship between the individual's motion trajectory vector and the safety guidance vector in the dynamic semantic model.
[0208] The fall detection module performs dense crowd risk monitoring in the evacuation bottleneck area identified by the dynamic semantic model, identifies targets that meet the fall pattern criteria as potential fall targets, and analyzes the displacement direction characteristics of the crowd around the potential fall targets.
[0209] The trampling warning module generates a trampling risk warning result when the displacement direction characteristics meet the centripetal convergence condition.
[0210] Each of the modules can be implemented in whole or in part through software, hardware, or a combination thereof. It supports hardware embedded in or independent of the processor in the computer device, and also supports software stored in the memory of the computer device, so that the processor can call and execute the operations corresponding to each of the above modules.
[0211] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario, characterized in that, include: S1. Perform coordinate transformation based on the calibration parameters of the surveillance camera to map the two-dimensional image coordinates to the top-view physical coordinate system, and fuse scene structure information to generate an initial scene geometric semantic model. S2. During the period when no fire alarm signal is received, perform object detection and contour extraction to obtain the physical contour of the object left in the monitoring screen. S3. Perform obstacle interference analysis, calculate the physical interference between the physical contour of the leftover object and the moving area of the fire door in the scene geometric semantic model. When the physical interference exceeds the preset interference threshold, it is determined that a passage blockage event has occurred. Based on the physical position of the leftover object, the scene geometric semantic model is dynamically updated to generate a dynamic semantic model. S4. In response to the received fire alarm signal, perform evacuation behavior recognition based on the dynamic semantic model and extract the individual movement trajectory vectors of the evacuating crowd; S5. Generate reverse behavior detection results by comparing the directional relationship between the individual motion trajectory vector and the safety guidance vector in the dynamic semantic model; S6. Perform dense crowd risk monitoring in the evacuation bottleneck area identified by the dynamic semantic model, identify targets that meet the fall pattern criteria as potential fall targets, and analyze the displacement direction characteristics of the crowd around the potential fall targets. S7. When the displacement direction characteristics meet the centripetal convergence condition, a stampede risk warning result is generated. The scene's geometric semantic model is dynamically updated based on the physical location of the remaining objects, generating a dynamic semantic model, including: S3.1 When a passage blockage event is determined to occur, the physical outline of the remaining object is marked as a dynamic obstacle area in the top-view physical coordinate system; S3.2 In the top-view physical coordinate system, path replanning is performed with the dynamic obstacle area as a constraint. The avoidance path from the upstream of the dynamic obstacle area to the safety exit is calculated, and the direction of the avoidance path at the key position is extracted and normalized to generate a dimensionless unit vector as the updated safety guidance vector. S3.3 Calculate the remaining passage width between the edge of the dynamic obstacle area and the wall boundaries on both sides of the passage. Mark the sections with remaining passage width less than the preset human passage width threshold as the updated evacuation bottleneck area. The dynamic obstacle area, the updated safety guidance vector, and the updated evacuation bottleneck area constitute the dynamic semantic model.
2. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 1, characterized in that, Based on the calibration parameters of the surveillance camera, a coordinate transformation is performed to map the two-dimensional image coordinates to the top-view physical coordinate system, and scene structure information is fused to generate an initial scene geometric semantic model, including: S1.1 Obtain the internal parameters of the surveillance camera and its external parameters relative to the ground, and calculate the homography transformation matrix from the image plane to the ground physical plane; S1.2 Identify the key structural points of the fire door in the image, and use the homography transformation matrix to transform the coordinates of the key structural points to the top-view physical coordinate system to obtain the physical position and physical dimensions of the door; S1.3 In the top-view physical coordinate system, based on the physical position and physical size of the door, construct a physical sweep area that represents the opening range of the door leaf, and combine it with the preset safety exit position to generate the initial safety guide vector, and mark the initial evacuation bottleneck area according to the physical width of the passage, together forming the initial scene geometric semantic model.
3. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 1, characterized in that, In the top-view physical coordinate system, path replanning is performed with dynamic obstacle areas as constraints. The avoidance path from the upstream of the dynamic obstacle area to the safety exit is calculated. Specifically, a graph search algorithm is used to discretize the top-view physical coordinate system into a grid, and the grid occupied by the dynamic obstacle area is set as an impassable node. The target passable path from the preset starting grid to the safety exit grid is searched as the avoidance path.
4. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 2, characterized in that, Extract the individual movement trajectory vectors of the evacuated crowd, including: S4.1 Perform target detection and tracking on the surveillance video sequence containing evacuated crowds to obtain the bounding rectangles of multiple human targets in consecutive image frames; S4.2 For each tracked human target, select the bottom feature points of its bounding rectangle, and use the homography transformation matrix to map the bottom feature points to the top-view physical coordinate system to obtain the sequence of physical location points of the human target. S4.
3. Based on the changes in the sequence of physical location points over consecutive timestamps, calculate the individual motion trajectory vector corresponding to each human target.
5. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 1, characterized in that, The reverse behavior detection results are generated by comparing the directional relationship between the individual motion trajectory vector and the safety guidance vector in the dynamic semantic model, including: S5.1 Obtain the updated safety guidance vector associated with the current individual's physical location in the dynamic semantic model; S5.2 Calculate the dot product of the individual motion trajectory vector and the updated safety guidance vector; S5.3 If the dot product value remains negative within a preset time duration threshold and the individual's motion trajectory vector crosses the predefined security alert boundary, then it is determined that the individual has committed a reverse intrusion behavior, and a reverse behavior detection result is generated.
6. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 1, characterized in that, Targets that meet the fall pattern criteria are identified as potential fall targets, including: S6.1 Within the evacuation bottleneck area, continuously monitor the shape and height parameters of the bounding rectangle of each human target. S6.2 When the morphological proportion parameters of a human target change from the set upright characteristic range to the fallen characteristic range within a preset time window, and the decrease in its position height parameter exceeds the preset height threshold, it is determined that the human target has fallen. S6.3 Mark the physical location of the target that has fallen as the potential center of mass of the fall, and identify the human target as a potential fall target.
7. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 1, characterized in that, Analyze the displacement direction characteristics of people around a potential fall victim, including: S6.
4. Delineate a circular monitoring area in the top-view physical coordinate system with the potential center of gravity of the fall as the center; S6.
5. Obtain other human targets that enter the circular monitoring area within the set time window, and calculate the individual motion trajectory vectors of these human targets; S6.6 Analyze the directionality of each individual's motion trajectory vector relative to the potential center of mass of the fall to obtain the displacement direction characteristics of the surrounding crowd.
8. The method for recognizing composite abnormal behaviors in a fire safety and security coupled scenario according to claim 7, characterized in that, When the displacement direction characteristics meet the centripetal convergence condition, a stampede risk warning result is generated, including: S7.1 Count the number of human targets around the potential center of mass of a fall whose individual motion trajectory vector continuously points to the center of mass within a continuous preset time frame. S7.2 If the number of human targets in the surrounding area exceeds the set threshold for the number of people to gather, then the displacement direction feature is determined to meet the centripetal gathering condition. S7.3 When there are potential falling targets and the centripetal convergence condition is met, a stampede risk warning result is generated.
9. A composite abnormal behavior recognition system in a fire protection and security coupled scenario, characterized in that, include: The semantic mapping module performs coordinate transformation based on the calibration parameters of the surveillance camera, maps the two-dimensional image coordinates to the top-view physical coordinate system, and integrates scene structure information to generate an initial scene geometric semantic model. The abandoned object detection module performs abandoned object detection and contour extraction during the period when no fire alarm signal is received, and obtains the physical contour of abandoned objects in the monitoring screen; The dynamic reconstruction module performs obstacle interference analysis, calculates the physical interference between the physical contour of the abandoned object and the moving area of the fire door in the scene geometric semantic model, determines that a passage blockage event has occurred when the physical interference exceeds the preset interference threshold, and dynamically updates the scene geometric semantic model based on the physical position of the abandoned object to generate a dynamic semantic model. The trajectory extraction module, in response to the received fire alarm signal, performs evacuation behavior recognition based on a dynamic semantic model and extracts the individual movement trajectory vectors of the evacuating crowd. The reverse driving detection module generates reverse driving behavior detection results by comparing the directional relationship between the individual's motion trajectory vector and the safety guidance vector in the dynamic semantic model. The fall detection module performs dense crowd risk monitoring in evacuation bottleneck areas identified by the dynamic semantic model, identifies targets that meet the fall pattern criteria as potential fall targets, and analyzes the displacement direction characteristics of people around potential fall targets. The stampede warning module generates a stampede risk warning result when the displacement direction characteristics meet the centripetal convergence condition. The scene's geometric semantic model is dynamically updated based on the physical location of the remaining objects, generating a dynamic semantic model, including: S3.1 When a passage blockage event is determined to occur, the physical outline of the remaining object is marked as a dynamic obstacle area in the top-view physical coordinate system; S3.2 In the top-view physical coordinate system, path replanning is performed with the dynamic obstacle area as a constraint. The avoidance path from the upstream of the dynamic obstacle area to the safety exit is calculated, and the direction of the avoidance path at the key position is extracted and normalized to generate a dimensionless unit vector as the updated safety guidance vector. S3.3 Calculate the remaining passage width between the edge of the dynamic obstacle area and the wall boundaries on both sides of the passage. Mark the sections with remaining passage width less than the preset human passage width threshold as the updated evacuation bottleneck area. The dynamic obstacle area, the updated safety guidance vector, and the updated evacuation bottleneck area constitute the dynamic semantic model.
Citation Information
Patent Citations
CN110807345A
CN121861567A