Fire Lane Occupancy Monitoring Method and System Based on Large Model
By combining multimodal large model and open target detection model, and adopting a static tracking mechanism and a modular system architecture, the existing fire channel occupation monitoring system has solved the problem of fixed target type, insufficient accuracy and poor scalability in detection targets, and achieved high intelligence, wide detection range and high accuracy of fire channel monitoring effects.
Patent Information
- Application Number
- CN202510348442.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing fire passage occupancy monitoring system has problems such as fixed detection target type, insufficient detection accuracy, high false alarm rate and poor system scalability, making it difficult to achieve monitoring effects with high intelligence, wide detection range and high accuracy.
A technical solution combining multimodal large model and open object detection model is adopted, and the monitoring images are periodically acquired and input multimodal large model for preliminary target recognition, dynamically output the set of labels to be detected, and combined with the static tracking mechanism and a modular system architecture, efficient monitoring and alarming of fire channels is achieved.
It significantly improves the intelligence level and detection capabilities of the system, expands the detection range, improves the detection accuracy, reduces false alarms, reduces system maintenance costs, and realizes all-weather automated monitoring.
Smart Images

Figure CN119851220B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology for safety monitoring, and particularly to a method and system for monitoring the occupancy of fire lanes based on large models. Background Art
[0002] With the continuous expansion of the urban scale and the continuous increase in population density, fire safety management faces increasingly severe challenges. As an important passage for evacuating people and fire trucks during a fire, the unobstructed state of the fire lane is directly related to the life and property safety of the people. However, due to problems such as parking difficulties, the phenomenon of fire lanes being occupied by motor vehicles, non-motor vehicles and other items is not uncommon, which seriously threatens public safety.
[0003] Currently, the monitoring of fire lane occupancy mainly adopts the following methods: one is to monitor through manual inspections, which has a high labor cost and cannot ensure 24-hour continuous monitoring; the other is to use traditional video monitoring and computer vision technology for monitoring. Although this method can achieve automated monitoring, it has the following deficiencies:
[0004] 1. The types of detected targets are fixed and difficult to expand. Traditional video analysis methods usually use specific open target detection models, which can only identify pre-defined fixed types of targets, such as vehicles, non-motor vehicles, etc. When new types of occupancy items appear, the model needs to be retrained, and the adaptability is poor.
[0005] 2. The detection accuracy is insufficient and the false alarm rate is high. Traditional computer vision models are prone to missed detections and false detections in complex scenarios, especially in cases of insufficient light, weather changes, etc. In addition, due to the lack of the ability to continuously track and analyze target objects, there are often repeated alarms or false alarms.
[0006] 3. The system scalability is poor. Existing systems often tightly couple open target detection models and business logic. When new monitoring objects need to be added or monitoring strategies need to be adjusted, the entire system needs to be massively modified, and the system maintenance cost is high.
[0007] Therefore, there is an urgent need for a fire lane occupancy monitoring method and system with higher intelligence, wider detection range and higher accuracy. With the development of artificial intelligence technology, especially the breakthrough progress made by multi-modal large models in the field of visual understanding, new technical means are provided to solve the above technical problems. Summary of the Invention
[0008] In order to solve the deficiencies existing in the above-mentioned prior art, the present invention aims to utilize the advantages of multi-modal large models and deep learning technology to achieve more intelligent and efficient fire lane occupancy monitoring.
[0009] To achieve the above-mentioned invention objective, the technical solution provided by the present invention includes:
[0010] A method for monitoring the occupancy of a fire passage based on a large model, including the steps of:
[0011] S1. Periodically obtain at least one frame of monitoring image of the target fire passage, input it into a multi-modal large model for preliminary target recognition, and use a preset prompt word template to guide the multi-modal large model to output a first set of tags to be detected for the monitoring image; the first set of tags to be detected includes Chinese and English tags;
[0012] S2. Use the first set of tags to be detected to update a second set of tags to be detected with a fixed length in chronological order. When the second set of tags to be detected changes, the AI engine obtains the video stream within the current period and decodes it to obtain several frames of video stream images, and sends the video stream images and capability parameters to an open target detection model;
[0013] S3. The open target detection model uses the capability parameters as input parameters to complete the detection of target objects in the video stream images and outputs a list of target object information;
[0014] S4. Perform static tracking on the target objects in the target fire passage area of interest according to the list of target object information and output a list of tracked target object information;
[0015] S5. When the tracked target objects in the list of tracked target object information meet the preset warning conditions, report a fire passage occupancy warning event.
[0016] Preferably, the capability parameters include: camera ID, video stream address, video stream format, area of interest, AI capability name, tracking duration threshold, interval warning period.
[0017] Preferably, the list of target object information includes: target location, target category, target confidence level, image timestamp.
[0018] Preferably, the method for static tracking in step S4 includes:
[0019] Filter the target objects in the list of target object information that do not belong to the target fire passage area of interest;
[0020] Merge the target objects with an intersection over union greater than a preset intersection threshold among the remaining target objects;
[0021] If the remaining target objects do not appear at the target location within a preset time, it is determined that the target object is lost.
[0022] Preferably, the list of tracked target object information includes: tracking ID, tracking status, target duration, target location, target category, target threshold, image timestamp.
[0023] Preferably, the warning conditions include:
[0024] The duration of the tracked target object in the target location in the list of tracked target object information is greater than a preset time;
[0025] The number of reports of the tracked target object in the list of tracked target object information within the interval warning period does not exceed 1 time.
[0026] Preferably, if the interval warning period is 0, the tracked target object in the list of tracked target object information reports only 1 time during the entire monitoring period.
[0027] The present invention also provides a fire passage occupancy monitoring system based on a large model, and the system is used to implement the above-mentioned fire passage occupancy monitoring method based on a large model.
[0028] Beneficial effects
[0029] 1. The present invention adopts a technical solution that combines a multi-modal large model and an open target detection model, significantly improving the intelligent level and detection ability of the system. Through a preset prompt word template, the multi-modal large model can automatically understand and analyze the monitoring scenario, and dynamically output a set of tags to be detected. This method breaks through the limitation that traditional systems can only recognize fixed types of targets, greatly expanding the detection range of the system. For example, when the fire passage is occupied by unconventional items such as temporarily stacked building materials and express packages, the system can still accurately identify and give an alarm in time.
[0030] 2. The present invention innovatively introduces a static tracking mechanism, effectively solving the problems of repeated alarms and false alarms in traditional systems. By continuously tracking and analyzing the target objects in the target fire passage area of interest, the system can accurately judge the duration of the occupancy behavior, and intelligently decide whether to trigger an alarm according to the preset warning conditions. This method greatly improves the detection accuracy of the system and reduces the occurrence of false alarms.
[0031] 3. The present invention adopts a modular and service-oriented system architecture design, significantly improving the scalability and maintainability of the system. By decoupling the multi-modal large model, the open target detection model and the static tracking module, the system can be flexibly expanded in function and optimized and upgraded. At the same time, the design of a standardized parameter transmission interface enables the system to conveniently access new AI capabilities, providing a good foundation for future function expansion.
[0032] 4. The present invention has significant social and economic values in practical applications. This system can achieve all-weather and automated monitoring of fire lanes, greatly reducing the cost of manual inspections. At the same time, the high accuracy rate and real-time alarm capabilities of the system effectively prevent the illegal act of fire lanes being occupied, improve the urban fire safety management level, and make important contributions to protecting the lives and property safety of the people. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic flowchart of a method for monitoring the occupation of fire lanes based on a large model provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0035] Embodiment 1
[0036] As Figure 1 shown, this embodiment discloses a method for monitoring the occupation of fire lanes based on a large model, including the steps of:
[0037] S1. Periodically obtain at least one frame of monitoring image of the target fire lane, input it into a multimodal large model for preliminary target recognition, and use a preset prompt word template to guide the multimodal large model to output a first set of tags to be detected for the monitoring image; the first set of tags to be detected includes Chinese and English tags.
[0038] In this step, first, at least one frame of monitoring image of the target fire lane is obtained through a camera at a preset time period (for example, at an interval of 1 minute). This periodic sampling method can not only ensure real-time monitoring of the status of the fire lane but also avoid waste of system resources caused by processing redundant image data.
[0039] After acquiring the monitoring image, it is input into the multimodal large model for preliminary target recognition. The multimodal large model can be any of the commonly used large models in the field. Preferably, the InternVL2 private deployment solution can be used, which adopts a dynamic resolution strategy to divide the image into blocks of different sizes, and adjusts the aspect ratio and resolution of the input image to better perceive the target object in the image. In order to guide the model to output standardized detection labels, the system uses a preset prompt word template.
[0040] Specifically, the structure of the prompt template can be divided into two roles: system / user. The specific examples are as follows:
[0041] system: "You are a visual assistant that can identify the target types on and around the fire escape in the image based on the image information input by the user, and provide a set of Chinese and English labels.\nThe target type name should be as simple as possible; for example, 'car' should be represented as 'car', 'motorcycle' as 'motorcycle', 'bicycle' as 'bicycle', and so on.\nDo not include special characters in the returned results, such as: \\n,..., etc. Multiple targets are separated by commas.\nPlease refer to the returned answer examples: Car, Motorcycle, Bicycle, Bicycle".
[0042] user: "Please identify the objects on and around the fire escape in the image and return a set of Chinese and English labels. Multiple objects should be separated by commas and no special characters should appear."
[0043] The multimodal large model will analyze the input monitoring image based on this prompt word template and output the first set of labels to be detected. This Chinese-English label design has the following advantages:
[0044] 1. Improve the versatility of the system and support applications in different language environments;
[0045] 2. Facilitate standardized docking with downstream open target detection models;
[0046] 3. It is conducive to system maintenance and data analysis;
[0047] It should be noted that the multimodal large model mainly plays the role of scene understanding and label generation in this step, providing input parameters for subsequent precise target detection. This design fully utilizes the advantages of the multimodal large model in scene understanding, while avoiding the performance overhead that may be caused by directly using the multimodal large model for target detection.
[0048] S2. Update the second set of tags to be detected with a fixed length in chronological order using the first set of tags to be detected. When the second set of tags to be detected changes, the AI engine acquires the video stream within the current cycle, decodes it to obtain several video stream images, and sends the video stream images and capability parameters to the open target detection model.
[0049] This step implements a dynamic update tag management mechanism and triggers the target detection process based on tag changes. This step can be elaborated in three main parts:
[0050] Firstly, it is the dynamic update mechanism of the tag set. The system maintains a second set of tags to be detected with a fixed length (e.g., 20), and this tag set is updated in a first-in, first-out (FIFO) manner. When a new first set of tags to be detected is obtained, the system adds the tags in it to the second set of tags to be detected in chronological order. If the length of the second set of tags to be detected exceeds 20, the earliest added tag will be automatically removed. This design ensures that the system always uses the latest detection targets and optimizes the use of system resources by limiting the number of tags.
[0051] Secondly, it is change detection and video stream processing. The system monitors the changes in the second set of tags to be detected. When a change is detected (such as adding a tag, deleting a tag, or modifying a tag), the AI engine will immediately start processing the video stream. Specifically, the AI engine acquires the video stream data within the current time cycle and converts the video stream into a series of image frames through a decoder. This decoding process takes into account parameters such as the resolution and frame rate of the video to ensure that the quality of the decoded images meets the requirements of target detection. It should be understood that the AI engine is an independent deep learning inference service system, which is responsible for performing actual visual recognition tasks. This service system has the following characteristics:
[0052] 1. Service architecture design: The AI engine adopts a service-oriented architecture and provides services externally through interfaces. This design enables decoupling between the main system and the AI inference engine, facilitating independent maintenance and upgrade.
[0053] 2. Parameter update mechanism: When the tag set changes, the reason why the system knows which detection capabilities need to be updated is that a mapping relationship management mechanism is implemented: the system maintains a correspondence table between tags and detection capabilities, and each tag is associated with specific detection capabilities.
[0054] 3. When the tag changes, the system can determine which detection capabilities need to be updated by querying this mapping relationship.
[0055] It should be understood that the AI engine can obtain only the video stream within the target period, or continuously obtain the video stream without interruption, and then transfer the video stream within the specified period to the open target detection model according to the period range. The specific acquisition method can be designed by those skilled in the art according to actual needs, and the present invention does not make further requirements.
[0056] In some preferred embodiments, an inference engine based on a deep learning framework, such as TensorRT and Pytorch, etc., can be used to implement the AI engine in the present invention. This part of the content is not the focus of the present invention, so it will not be elaborated here. Those skilled in the art can make a choice in the prior art according to actual needs.
[0057] Finally, it is the parameter distribution link. The decoded video stream image is distributed to the open target detection model together with the complete capability parameters. The capability parameters are a set of configuration information, which define the working mode and boundary conditions of the open target detection model when performing the detection task. In the present invention, the capability parameters include the following core contents: Basic configuration: including camera ID, video stream address and format (RTSP), etc. These parameters tell the model where to obtain the input data; Detection range: The model is limited to only focus on the specific area where the fire passage is located through the region of interest parameter; Target type: The object categories to be detected are specified through the second set of tags to be detected; Time control: including the tracking duration threshold and the interval alarm period, which are used to control the time dimension requirements of the detection.
[0058] The relationship between these capability parameters and the open target detection model is a "configuration and execution" relationship. The open target detection model is a deep learning model that actually executes the object recognition task, while the capability parameters tell the model what to detect, where to detect, and how to detect. It should be understood that the open target detection model described in the present invention is a type of deep learning model dedicated to object localization and recognition in images. It includes both the basic functions of traditional target detection and breaks through the limitation that traditional models can only recognize fixed categories. Represented by advanced models such as DINO (DETR with Improved deNoising anchOr boxes). The open feature of this type of model makes it have stronger adaptability and practical value in the actual applications within the scope involved in the present invention, and can better meet the target detection requirements in various scenarios. Those skilled in the art can make a reasonable choice in the prior art according to the on-site hardware resource conditions and the requirements of real-time performance and detection accuracy, and the present invention does not make further limitations.
[0059] In some preferred embodiments, specific examples of the capability parameters are given. These capability parameters are a parameter set containing multiple configuration items, specifically including:
[0060] Camera ID: Used to uniquely identify the video input source;
[0061] Video stream address: Specifies the location of the data source;
[0062] Video stream format: Specifies the video format, such as RTSP format;
[0063] Region of interest: Limits the detection range in the fire lane;
[0064] AI capability name: Set to FireLaneObstruction;
[0065] Tracking duration threshold: Sets the time requirement for target tracking;
[0066] Interval alarm period: Controls the triggering frequency of the alarm.
[0067] S3. The open target detection model takes the capability parameters as input parameters, completes the detection of target objects in the video stream image, and outputs a list of target object information.
[0068] The open target detection model will parse the received capability parameters. These parameters affect multiple aspects of the detection process. For example, the camera ID and video stream address determine the data source, and the region of interest parameter defines the image range that needs to be concerned about. These parameters together construct a complete context for a detection task. The list of target object information contains detailed information about each detected target object. Preferably, it includes the target location, target category, target confidence, and image timestamp.
[0069] S4. Perform static tracking on the target objects in the target fire lane region of interest based on the list of target object information, and output a list of tracked target object information.
[0070] The static tracking is a method for analyzing the persistence of objects based on the target detection results. Different from traditional video object tracking algorithms (such as KCF, SORT, etc.), static tracking mainly focuses on the continuous presence state of target objects in a specific region, rather than precisely tracking the movement trajectory of the target. In some preferred embodiments, specific methods for implementing static tracking are given, including:
[0071] First, perform region filtering on the list of target object information. Specifically, check whether the position coordinates of each target object are within the predefined fire lane region of interest.
[0072] Secondly, perform target merging processing. When multiple similar targets are detected (such as the same vehicle being detected repeatedly), calculate their intersection over union (IOU), and merge the targets with an IOU exceeding a preset threshold (for example, 0.5) into the same tracked target.
[0073] Then, maintain a status record for each target passed the screening, and generate a list of tracking target object information. This list contains richer information. In some preferred embodiments, the list of tracking target object information includes: tracking ID, tracking status, target duration, target location, target category, target threshold, image timestamp.
[0074] Finally, perform target loss judgment. If there is no new detection result for a certain tracking target within a preset time (e.g., 5 seconds), mark it as "lost" status.
[0075] The above static tracking mechanism has low computational overhead and does not require complex motion prediction; it has good fault tolerance for short-term occlusion of targets; it can accurately count the continuous stay time of targets; and it avoids repeated counting caused by short-term movement of targets.
[0076] S5. When the tracking target object in the list of tracking target object information meets the preset alarm conditions, report the fire lane occupancy alarm event. The alarm mechanism is the final decision-making link of the system. It judges whether to trigger an alarm by evaluating the status of the tracking target. Its design can be carried out by those skilled in the art according to the actual needs of the site. In some preferred embodiments, a method for reporting the fire lane occupancy alarm event is given by checking whether each tracking target meets two key alarm conditions, namely duration and alarm frequency control, specifically including:
[0077] Duration judgment: Check whether the stay time of the target in the fire lane area exceeds the preset tracking duration threshold.
[0078] Alarm frequency control: Avoid repeated alarms by means of an interval alarm period parameter to ensure that the same target only triggers one alarm within a specified time period.
[0079] When a tracking target meets both of these conditions at the same time, the system will generate an alarm event. The data structure of the alarm event contains complete on-site information.
[0080] The design of this alarm mechanism takes into account the requirements of multiple actual application scenarios:
[0081] 1. Through duration judgment, false alarms for temporarily passing objects are avoided. For example, a short stay of a vehicle for temporary parking and loading / unloading goods will not trigger an alarm.
[0082] 2. Through alarm frequency control, the system is prevented from generating junk alarms. For example, even if a vehicle continuously parks in the fire lane, the system will not repeatedly report alarms frequently.
[0083] 3. The complete alarm information facilitates subsequent manual processing and data analysis. On-site managers can quickly locate problems based on the alarm information, while data analysts can use this data for statistical analysis and trend research.
[0084] In the actual scenario of fire lane monitoring, some long-term occupation situations may be encountered. For example, a certain vehicle owner may continuously park their vehicle in the fire lane for several days or even weeks. In this case, if alarms are continuously sent according to the regular interval alarm cycle (such as every 4 hours), a large number of duplicate alarm messages will be generated. These duplicate alarms will not only occupy system resources, but more importantly, will lead to problems such as information redundancy and reduced processing efficiency. Therefore, in some preferred embodiments, by setting the interval alarm cycle to the special value of 0, a "one-time alarm" mechanism is provided. This means that for the same occupied object, the system will only send an alarm once when it is first detected to meet the alarm conditions, and no further duplicate alarms will be sent.
[0085] In some other preferred embodiments, the present invention also provides a fire lane occupancy monitoring system based on a large model, and the system is used for the above-mentioned fire lane occupancy monitoring method based on a large model.
[0086] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A fire passage occupancy monitoring method based on a large model, characterized in that: Includes steps: S1. Periodically acquire at least one frame of monitoring image of the target fire passage, input it into the multimodal large model for preliminary target recognition, and use a preset prompt word template to guide the multimodal large model to output a first set of labels to be detected for the monitoring image; the first set of labels to be detected includes Chinese and English labels; S2. Use the first set of tags to be detected to update the second set of tags to be detected with a fixed length in chronological order. When the second set of tags to be detected changes, the AI engine obtains the video stream in the current period and decodes it to obtain several frames of video stream images, and sends the video stream images and capability parameters to the open target detection model; S3. The open target detection model takes the capability parameter as an input parameter, completes the target object detection of the video stream image, and outputs a target object information list; S4. Static tracking of the target object in the target fire channel area of interest is performed according to the target object information list, and the tracking target object information list is output; S5. When the tracking target object in the tracking target object information list meets the preset alarm conditions, the fire channel occupancy alarm event is reported; The target object information list includes: target location, target category, target confidence, and image timestamp; The method for performing static tracking in step S4 includes: Filter the target objects in the target object information list that do not belong to the target fire channel interest area; Merge target objects whose intersection-and-union ratio is greater than a preset intersection-and-union threshold among the remaining target objects; If the remaining target objects do not appear at the target location within the preset time, it is determined that the target object is lost.
2. The fire passage occupancy monitoring method based on a large model as claimed in claim 1 is characterized in that: The capability parameters include: camera ID, video stream address, video stream format, area of interest, AI capability name, tracking time threshold, and interval alarm period.
3. The fire passage occupancy monitoring method based on a large model as claimed in claim 1 is characterized in that: The tracking target object information list includes: tracking ID, tracking status, target duration, target position, target category, target threshold, and image timestamp.
4. The fire passage occupancy monitoring method based on a large model as claimed in claim 2 is characterized in that: The alarm conditions include: The duration of the tracking target object in the tracking target object information list at the target location is greater than the preset time; The number of reports of the tracking target objects in the tracking target object information list within the interval alarm period shall not exceed 1 time.
5. The fire passage occupancy monitoring method based on a large model as claimed in claim 4 is characterized in that: If the interval alarm period is 0, the tracking target object in the tracking target object information list is reported only once in the whole monitoring period.
6. Fire passage occupancy monitoring system based on large model, characterized by: The system is used to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Target tracking method, device and equipment based on AI visual identification
CN119649063A
AI business management system based on enterprise large model data
CN119671221A