Interactive System for Monitoring Occupation of Fire Escape Routes

By combining multimodal large model and open object detection model, a dual identification mechanism is built, which solves the problems of low detection accuracy and simple interaction interface in fire passage occupation monitoring technology, and efficient and intelligent monitoring and interaction are achieved, significantly improving the intelligence level and accuracy of the system.

CN119888626BActive Publication Date: 2025-06-03CHENGDU KOALA URAN TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510345195.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-03
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing fire passage occupancy monitoring technology has problems such as low detection accuracy, inability to identify diverse occupancy scenarios, simple interaction interface, single alarm information display methods, and lack of effective information feedback mechanisms, making it difficult to achieve efficient and intelligent monitoring and interaction.

Method used

A technology combining multimodal large model and open object detection model is adopted to build a dual recognition mechanism to achieve preliminary and accurate target recognition of fire channel monitoring images. The system includes a tag acquisition module, a built-in AI engine, an open object detection model and an alarm module. Through static tracking and preset alarm conditions, real-time visual display, intelligent alarm and interactive handling are achieved.

Benefits of technology

It significantly improves the intelligence level and accuracy of fire passage occupation monitoring, realizes flexible and efficient target tracking, provides rich interactive functions and intuitive visual interface, reduces the probability of false alarms and missed alarms, and improves the ease of use and management efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888626B_ABST
    Figure CN119888626B_ABST
Patent Text Reader

Abstract

The present invention discloses an interaction system for monitoring the occupation of fire lanes, comprising: an image acquisition device, a display device, and a server; the server includes: a label acquisition module for acquiring monitoring images, inputting them into a multi-modal large model for preliminary target recognition to output a first set of labels to be detected, and updating a second set of labels to be detected with a fixed length, and when the second set of labels to be detected changes, notifying the built-in AI engine to start; a built-in AI engine for acquiring the video stream within the current period and sending the video stream images and capability parameters to an open target detection model; an open target detection model for using the capability parameters as input parameters to complete the detection of target objects in the video stream images and output a list of target object information; a module for statically tracking the target objects according to the list of target object information; and an alarm module for reporting a fire lane occupation alarm event to the user through the display device when a preset alarm condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of safety protection and detection, and particularly to an interactive system for monitoring the occupation of fire corridors. Background Art

[0002] Fire corridors are important passages in buildings for emergency evacuation in case of fire and fire fighting and rescue. Keeping fire corridors unobstructed is of great significance for ensuring the safety of people's lives. However, in actual applications, fire corridors are often randomly occupied, such as illegally parking vehicles, piling up sundries, etc. This situation seriously endangers the fire safety of buildings.

[0003] Currently, the monitoring and interaction of the occupation of fire corridors mainly adopt the following methods:

[0004] 1. Manual inspection method. This method requires security personnel to regularly patrol and inspect the fire corridors and report the occupation situation to the management center through simple interaction methods such as walkie-talkies or telephones. However, this method has problems such as high labor costs, low monitoring efficiency, untimely information transmission, and inability to ensure round-the-clock monitoring.

[0005] 2. Traditional video monitoring method. By installing ordinary monitoring cameras in the fire corridors, the duty personnel can observe in real time whether there is occupation through the display screen in the monitoring center and interact with the on-site personnel through the intercom system. Although this method realizes basic visual interaction, the interaction method is single, and it is easy to miss the occupation situation due to the distraction of personnel's attention.

[0006] 3. Automatic monitoring method based on simple image recognition. This method uses computer vision technology to analyze the monitoring video to detect whether the fire corridor is occupied. However, the existing technology has the following deficiencies: the detection accuracy is relatively low, and false alarms and missed alarms are prone to occur; it is unable to accurately identify diverse occupation scenarios; the interaction interface is simple and crude, lacking intuitive interaction functions such as real-time video stream display and target annotation; the display method of alarm information is single and cannot meet the interaction needs in different scenarios; there is no effective information feedback mechanism between the system and the user, which affects the disposal efficiency.

[0007] With the development of artificial intelligence technology, especially the progress of multi-modal large models and object detection technology, new technical means have been provided to solve the above problems. However, at present, no fire corridor occupation monitoring solution that combines multi-modal large models with open object detection models and provides rich interaction functions has been seen. Summary of the Invention

[0008] In order to solve the above technical problems, the present invention aims to provide an intelligent monitoring system with a good human-computer interaction experience, realizing real-time visual display, intelligent alarm and interactive disposal of occupation behaviors.

[0009] To achieve the above-mentioned invention objectives, the technical solutions provided by the present invention include:

[0010] An interaction system for monitoring the occupation of a fire passage, comprising:

[0011] An image acquisition device for acquiring monitoring images of a target fire passage;

[0012] A display device for interacting with users;

[0013] A server respectively signal-connected to the image acquisition device and the display device; the server is configured to: based on the monitoring images of the target fire passage acquired by the image acquisition device, determine whether the fire passage is currently occupied, and if so, alarm the user through the display device;

[0014] The server includes a label acquisition module, a built-in AI engine, an open target detection model, and an alarm module;

[0015] The label acquisition module is configured to be respectively signal-connected to a multimodal large model and the image acquisition device, and is used to periodically acquire at least one frame of monitoring image of the target fire passage, input it into the multimodal large model for preliminary target recognition, and use a preset prompt word template to guide the multimodal large model to output a first set of to-be-detected labels of the monitoring image; use the first set of to-be-detected labels to update a second set of to-be-detected labels with a fixed length in chronological order, and when the second set of to-be-detected labels changes, notify the built-in AI engine to start; the first set of to-be-detected labels includes Chinese and English labels;

[0016] The built-in AI engine is signal-connected to the image acquisition device and the open target detection model, and is used to acquire the video stream in the current period and decode it to obtain several video stream images, and send the video stream images and capability parameters to the open target detection model;

[0017] The open target detection model is signal-connected to the alarm module, and is used to use the capability parameters as input parameters to complete the target object detection of the video stream images and output a list of target object information;

[0018] The alarm module is signal-connected to the display device, and is used to perform static tracking on the target objects in the target fire passage interest area according to the list of target object information and output a list of tracked target object information; when the tracked target objects in the list of tracked target object information meet the preset alarm conditions, report the fire passage occupation alarm event to the user through the display device.

[0019] Preferably, the alarm module includes a target tracking unit and an event reporting unit;

[0020] The target tracking unit is used to perform static tracking on the target objects in the target fire corridor area of interest according to the target object information list, and output a list of tracked target object information;

[0021] The event reporting unit is respectively connected to the target tracking unit and the display device by signals, and is used to judge whether the tracked target objects in the list of tracked target object information meet the preset warning conditions. If so, it reports warning information to the display device.

[0022] Preferably, the warning module includes a video fusion unit and a video display unit;

[0023] The video fusion unit is respectively connected to the built-in AI engine and the open target detection model by signals, and is used to obtain video stream images and a list of target object information, and output a real-time video stream and target object recognition frames to the video display unit;

[0024] The video display unit is connected to the display device by signals, and is used to obtain the real-time video stream and target object recognition frames, and after rendering, send them to the display device for presenting to the user.

[0025] Preferably, the capability parameters include: camera ID, video stream address, video stream format, area of interest, AI capability name, tracking duration threshold, interval warning period.

[0026] Preferably, the target object information list includes: target location, target category, target confidence, image timestamp; the tracked target object information list includes: tracking ID, tracking status, target duration, target location, target category, target threshold, image timestamp.

[0027] Preferably, the method for the target tracking unit to perform static tracking includes:

[0028] Filter the target objects in the target object information list that do not belong to the target fire corridor area of interest;

[0029] Merge the target objects with an intersection over union greater than the preset intersection threshold among the remaining target objects;

[0030] If the remaining target objects do not appear at the target location within the preset time, it is determined that the target object is lost.

[0031] Preferably, the warning conditions include:

[0032] The duration of the tracked target object in the list of tracked target object information at the target location is greater than the preset time;

[0033] The number of reports of the tracked target object in the list of tracked target object information within the interval warning period does not exceed 1 time.

[0034] Preferably, if the interval warning period is 0, the tracked target objects in the tracked target object information list are reported only once during the entire monitoring period.

[0035] Preferably, the method for aligning the target object recognition frame with the target object in the real-time video stream when the video display unit finishes rendering includes:

[0036] Obtain the target object information list in the real-time video stream, query the image timestamp, match and obtain the target object information list and the real-time video stream data packet with the same timestamp, and send it to the display device for display to the user after rendering.

[0037] Preferably, if the frame rate of the target object information list with the same timestamp is lower than the frame rate of the real-time video stream, cache the target information list data and the real-time video stream data packet within the preset time threshold range, and use the method of linear interpolation to supplement the frame rate for the cached target information list data.

[0038] Beneficial effects

[0039] 1. Improved the intelligence level and accuracy of fire lane occupancy monitoring: This application combines a multi-modal large model with an open all-object detection model to construct a dual recognition mechanism: First, use the multi-modal large model to perform preliminary recognition on the monitoring image and output the set of tags to be detected; then use a professional object detection model for accurate recognition. This innovative technical solution significantly improves the detection accuracy of the system and effectively reduces the probability of false alarms and missed alarms.

[0040] 2. Achieved more flexible and efficient target tracking: By setting up a tag acquisition module, the system can periodically obtain monitoring images and dynamically update the set of tags to be detected, enabling the object detection model to adapt to scene changes in a timely manner. At the same time, through the static tracking mechanism and the preset intersection over union threshold, the system can effectively identify and track continuously occupied target objects, avoid repeated alarms, and improve the operating efficiency of the system.

[0041] 3. Provided rich interaction functions and an intuitive visualization interface: The system of this application combines the real-time video stream with the target object recognition frame through the video fusion unit and the video display unit for fusion display, enabling users to intuitively understand the monitoring results. The system supports real-time display of the detection and tracking status, and pushes alarm information to users in a timely manner through the display device, greatly enhancing the interaction experience and usage effect of the system.

[0042] 4. Avoid the phenomenon of floating frames: The system of this application aligns the target object information list with the time stamps of the real-time video stream data packets, and fills in the frame rate through linear interpolation when the frame rates do not match, so as to eliminate the phenomenon of out-of-sync between the target frame and the picture when presenting to the user, thereby improving the user interaction experience. 5. Reduce the system maintenance and usage costs: The system of this application adopts an automated monitoring and warning mechanism, which does not require continuous manual monitoring, significantly reducing the labor cost. The intelligent warning mechanism of the system can accurately locate the occupied position, improving the problem handling efficiency. Through an intuitive interaction interface, the user training cost is reduced, and the usability of the system is improved. Brief Description of the Drawings

[0043] Figure 1 It is a schematic structural diagram of an interactive system for fire passage occupancy monitoring provided in a preferred embodiment of the present invention. Detailed Embodiment

[0044] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0045] As Figure 1 shown, this embodiment provides an interactive system for fire passage occupancy monitoring, including:

[0046] An image acquisition device for acquiring monitoring images of the target fire passage.

[0047] A display device for interacting with users. In this embodiment, the display device is an important component for realizing information interaction between the system and users, and devices with display functions such as monitoring display screens in a fire control room, mobile terminal devices, or computer monitors can be used. The display device provides users with comprehensive monitoring information and an interaction interface by dividing into functional areas such as real-time monitoring, target recognition display, alarm information display, system status display, and parameter configuration. Among them, the real-time monitoring area displays the processed video stream, including the recognition frame and tracking trajectory of the target object; the alarm information area displays the detailed information of alarm events in real time and attracts the user's attention through sound and light prompts; the system status area feeds back the working status of each functional module, such as the connection status of the image acquisition device, the running status of the model, etc. The display device also supports users to perform interactive functions such as adjusting monitoring parameters, viewing historical alarm records, confirming and processing alarm events through touch operations, thereby realizing efficient information interaction between the system and users. This intuitive and convenient interaction method enables users to timely grasp the occupancy situation of the fire passage, significantly improving the usability and management efficiency of the system. It should be noted that those skilled in the art can appropriately adjust and expand the functions and interaction methods of the display device according to actual needs, and these changes should fall within the protection scope of this application.

[0048] A server respectively signal-connected to an image acquisition device and a display device; the server is configured to: judge whether the target fire passage is currently occupied according to the monitoring image of the target fire passage acquired by the image acquisition device, and if so, alarm to the user through the display device. In this embodiment, the server serves as the core processing unit of the system and establishes a signal connection with the image acquisition device and the display device through a network or other communication means. The server receives the monitoring images of the target fire passage from the image acquisition device, and these images can be single-frame images or continuous video streams. The server performs intelligent analysis and processing on the received monitoring images, and judges whether the fire passage is occupied through the built-in human algorithm, such as whether there are vehicles parked illegally or whether there are sundries stacked. When the server detects that the fire passage is occupied, it will immediately trigger an alarm mechanism and send alarm information to the display device through the signal connection channel with the display device. The alarm information may include key information such as the occupancy location, the type of occupancy object, and the occurrence time. After receiving these information, the display device will notify the user in a preset manner (such as screen prompts, sound reminders, etc.) in a timely manner, enabling the user to quickly understand the abnormal situation of the fire passage and take corresponding treatment measures.

[0049] The server includes a tag acquisition module, a built-in AI engine, an open target detection model, and an alarm module.

[0050] The label acquisition module is configured to be signal-connected to the multi-modal large model and the image acquisition device respectively, and is used to periodically acquire at least one frame of monitoring image of the target fire passage, input it into the multi-modal large model for preliminary target recognition, and use a preset prompt word template to guide the multi-modal large model to output the first set of labels to be detected for the monitoring image; update the second set of labels to be detected with a fixed length in chronological order using the first set of labels to be detected, and when the second set of labels to be detected changes, notify the built-in AI engine to start; the first set of labels to be detected includes Chinese and English labels.

[0051] Among them, at least one frame of monitoring image of the target fire passage is acquired by the image acquisition device at a preset time period (for example, every 1 minute). This periodic sampling method can not only ensure the real-time monitoring of the state of the fire passage, but also avoid the waste of system resources caused by processing redundant image data.

[0052] After the monitoring image is acquired, it is input into the multi-modal large model for preliminary target recognition. The multi-modal large model can be any one of the commonly used large models in the field. Preferably, the InternVL2 private deployment solution can be adopted. This solution supports the dynamic resolution strategy, divides the image into blocks of different sizes, and adjusts based on the aspect ratio and resolution of the input image, which can better perceive the target objects in the image. In order to guide the model to output standardized detection labels, the system adopts a preset prompt word template.

[0053] Specifically, the structure of the prompt word template can be divided into two roles: system / user. The specific examples are as follows:

[0054] system: "You are a visual assistant who can identify the target types on and around the fire passage in the picture according to the picture information input by the user, and provide a set of Chinese and English labels. \nThe name of the target type should be answered as simply as possible; for example, represent 'car' as 'car','motorcycle' as'motorcycle', 'bicycle' as 'bicycle', and so on. \nDo not include special characters in the returned result, such as: \n,..., etc. Multiple targets are separated by commas. \nPlease refer to the returned answer example: car,Car,motorcycle,Motorcycle,bicycle,Bicycle".

[0055] user: "Please identify the targets on and around the fire passage in the picture, and return a set of Chinese and English labels, separated by commas for multiple targets, and do not appear special characters".

[0056] Based on this prompt word template, the multi-modal large model will analyze the input monitoring image and output the first set of labels to be detected. This Chinese-English bilingual label design has the following advantages:

[0057] 1. It improves the universality of the system and supports applications in different language environments;

[0058] 2. It is convenient for standardized docking with downstream open target detection models;

[0059] 3. It is beneficial to the maintenance of the system and data analysis;

[0060] It should be noted that in this step, the multi-modal large model mainly plays the role of scene understanding and label generation, providing input parameters for subsequent precise target detection. This design makes full use of the advantages of the multi-modal large model in scene understanding, while avoiding the possible performance overhead of directly using the multi-modal large model for target detection.

[0061] Furthermore, the label acquisition module also implements a dynamic update mechanism for the label set. Specifically, the system maintains a second label set to be detected with a fixed length (e.g., 20), and this label set is updated in a first-in-first-out (FIFO) manner. When a new first label set to be detected is obtained, the system will add the labels in it to the second label set to be detected in chronological order. If the length of the second label set to be detected exceeds 20, the earliest added label will be automatically removed. This design ensures that the system always uses the latest detection targets and optimizes the use of system resources by limiting the number of labels.

[0062] Secondly, it is change detection and video stream processing. The system will monitor the changes in the second label set to be detected. When a change is detected (such as adding a label, deleting a label, or modifying a label), it will notify the built-in AI engine to start and immediately begin to process the video stream. Specifically, the AI engine will obtain the video stream data within the current time period and convert the video stream into a series of image frames through a decoder. This decoding process will consider parameters such as the resolution and frame rate of the video to ensure that the quality of the decoded images meets the requirements of target detection. It should be understood that the AI engine is an independent deep learning inference service system, which is responsible for performing actual visual recognition tasks. This service system has the following characteristics:

[0063] 1. Service architecture design: The AI engine adopts a service-oriented architecture and provides services externally through interfaces. This design enables decoupling of the main system and the AI inference engine, facilitating independent maintenance and upgrade.

[0064] 2. Parameter update mechanism: When the label set changes, the system knows which detection capabilities need to be updated because it has implemented a mapping relationship management mechanism: The system maintains a correspondence table between labels and detection capabilities, and each label is associated with specific detection capabilities.

[0065] 3. When the label changes, the system can determine which detection capabilities need to be updated by querying this mapping relationship.

[0066] It should be understood that the AI engine can obtain only the video stream within the target period, or continuously obtain the video stream without interruption, and then transfer the video stream within the specified period to the open target detection model according to the period range. The specific acquisition method can be designed by those skilled in the art according to actual needs, and the present invention does not make further requirements.

[0067] In some preferred embodiments, an inference engine based on a deep learning framework, such as TensorRT and Pytorch, etc., can be used to implement the AI engine in the present invention. This part of the content is not the focus of the present invention, so it will not be elaborated here. Those skilled in the art can make selections from the existing technologies according to actual needs.

[0068] The built-in AI engine is signal-connected to the image acquisition device and the open target detection model, and is used to obtain the video stream within the current period, decode it to obtain several video stream images, and send the video stream images and capability parameters to the open target detection model. Specifically, the built-in AI engine sends the decoded video stream images together with the complete capability parameters to the open target detection model. The capability parameters are a set of configuration information that defines the working mode and boundary conditions of the open target detection model when performing the detection task. In the present invention, the capability parameters include the following core contents: Basic configuration: including camera ID, video stream address and format (RTSP), etc. These parameters tell the model where to obtain the input data; Detection range: The model only needs to focus on a specific area where the fire passage is located by defining the region of interest parameter; Target type: The object categories to be detected are specified through the second set of labels to be detected; Time control: including the tracking duration threshold and the interval warning period, which are used to control the time dimension requirements of the detection.

[0069] The relationship between these capability parameters and the open object detection model is a "configuration and execution" relationship. The open object detection model is a deep learning model that actually executes object recognition tasks, while the capability parameters tell the model "what to detect", "where to detect", and "how to detect". It should be understood that the open object detection model described in the present invention is a type of deep learning model specifically used for object localization and recognition in images. It not only includes the basic functions of traditional object detection but also breaks through the limitation that traditional models can only recognize fixed categories. Represented by advanced models such as DINO (DETR with Improved deNoising ObjectDetector). The open feature of this type of model makes it have stronger adaptability and practical value in practical applications within the scope involved in the present invention, and can better meet the object detection requirements in various scenarios. Those skilled in the art can make reasonable selections in the prior art according to the on-site hardware resource conditions and the requirements of real-time performance and detection accuracy, and the present invention does not make further limitations.

[0070] In some preferred embodiments, specific examples of the capability parameters are given. These capability parameters are a parameter set containing multiple configuration items, specifically including:

[0071] Camera ID: Used to uniquely identify the video input source;

[0072] Video stream address: Specify the data source location;

[0073] Video stream format: Specify the video format, such as RTSP format;

[0074] Region of interest: Limit the detection range in the fire lane;

[0075] AI capability name: Set to FireLaneObstruction;

[0076] Tracking duration threshold: Set the time requirement for target tracking;

[0077] Interval alarm period: Control the triggering frequency of the alarm.

[0078] The open object detection model is signal-connected to the alarm module, and is used to take the capability parameters as input parameters to complete the detection of target objects in the video stream image and output a list of target object information. The open object detection model will parse the received capability parameters. These parameters will affect multiple aspects of the detection process. For example, the camera ID and the video stream address determine the data source, and the region of interest parameter defines the image range that needs to be concerned about. These parameters jointly construct a complete context of a detection task. The list of target object information contains detailed information of each detected target object. Preferably, it includes the target position, target category, target confidence, and image timestamp.

[0079] The alarm module is signal-connected to the display device and is used to perform static tracking on the target objects in the target fire passage area of interest according to the target object information list, and output a list of tracked target object information; when the tracked target objects in the list of tracked target object information meet the preset alarm conditions, a fire passage occupancy alarm event is reported to the user through the display device.

[0080] The alarm mechanism is the final decision-making link of the system. It judges whether an alarm needs to be triggered by evaluating the state of the tracked target. Its design can be carried out by those skilled in the art according to the actual needs of the site. In some preferred embodiments, the alarm module includes a target tracking unit and an event reporting unit;

[0081] The target tracking unit is used to perform static tracking on the target objects in the target fire passage area of interest according to the target object information list, and output a list of tracked target object information; the static tracking is an object persistence analysis method based on the target detection result. Different from traditional video target tracking algorithms (such as KCF, SORT, etc.), static tracking mainly focuses on the continuous presence state of the target object in a specific area, and does not require accurate tracking of the target's motion trajectory. In some preferred embodiments, specific methods for implementing static tracking are given, including:

[0082] First, perform area filtering on the target object information list. Specifically, check whether the position coordinates of each target object are within the predefined fire passage area of interest.

[0083] Secondly, perform target merging processing. When multiple similar targets are detected (for example, the same vehicle is detected repeatedly), calculate their intersection over union (IOU), and merge the targets with an IOU exceeding the preset threshold (for example, 0.5) into the same tracked target.

[0084] Then, maintain a status record for each target passing the screening, and generate a list of tracked target object information. This list contains richer information. In some preferred embodiments, the list of tracked target object information includes: tracking ID, tracking status, target duration, target location, target category, target threshold, image timestamp.

[0085] Finally, perform target loss judgment. If a certain tracked target has no new detection results within a preset time (for example, 5 seconds), mark it as the "lost" state.

[0086] The above static tracking mechanism has low computational overhead, does not require complex motion prediction; has good fault tolerance for short-term occlusion of the target; can accurately count the continuous stay time of the target; and avoids repeated counting caused by short-term movement of the target.

[0087] The event reporting unit is respectively connected to the target tracking unit and the display device in terms of signals, and is used to determine whether the tracked target objects in the list of tracked target object information meet the preset warning conditions. If so, it reports the warning information to the display device.

[0088] In some preferred embodiments, a method for reporting the warning event of fire lane occupation is provided by checking whether each tracked target meets two key warning conditions, namely duration and warning frequency control. Specifically, it includes:

[0089] Duration judgment: Check whether the staying time of the target in the fire lane area exceeds the preset tracking duration threshold.

[0090] Warning frequency control: Avoid repeated warnings by means of the interval warning cycle parameter, ensuring that the same target only triggers one warning within the specified time period.

[0091] When a tracked target meets both of these conditions at the same time, the system will generate a warning event. The data structure of the warning event contains complete on-site information.

[0092] The design of this warning mechanism takes into account the requirements of multiple actual application scenarios:

[0093] 1. Through duration judgment, false alarms for temporarily passing objects are avoided. For example, a short stay of a vehicle for temporary parking and loading / unloading goods will not trigger a warning.

[0094] 2. Through warning frequency control, the system is prevented from generating junk warnings. For example, even if a vehicle continuously parks in the fire lane, the system will not repeatedly report warnings frequently.

[0095] 3. The complete warning information is convenient for subsequent manual processing and data analysis. On-site management personnel can quickly locate problems based on the warning information, while data analysis personnel can use these data for statistical analysis and trend research.

[0096] In the actual scenario of fire lane monitoring, some long-term occupation situations may be encountered. For example, a certain vehicle owner may continuously park the vehicle in the fire lane for several days or even weeks. In this case, if warning messages are continuously sent according to the regular interval warning cycle (such as every 4 hours), a large number of repeated warning messages will be generated. These repeated warnings will not only occupy system resources, but more importantly, will lead to problems such as information redundancy and reduced processing efficiency. Therefore, in some preferred embodiments, by setting the interval warning cycle to the special value of 0, a "one-time warning" mechanism is provided. This means that for the same occupied object, the system will only send one warning when it is first detected that it meets the warning conditions, and no repeated warnings will be sent afterwards.

[0097] In order to achieve real-time visual display of the detection results of fire passage occupancy and improve the human-computer interaction effect of the system, in some preferred embodiments, the warning module includes a video fusion unit and a video display unit. The video fusion unit is respectively connected to the built-in AI engine and the open target detection model in a signal connection, and is used to obtain video stream images and a list of target object information, and output a real-time video stream and target object recognition frames to the video display unit; the video display unit is connected to the display device in a signal connection, and is used to obtain the real-time video stream and target object recognition frames, and send them to the display device for rendering and then presenting to the user after rendering.

[0098] The design of this warning module enables the system to fuse the video stream images from the built-in AI engine and the target object information output by the open target detection model in real time, accurately superimpose the target recognition frames on the original video screen, and ensure the display effect through special rendering processing. This modular design not only solves the problem of the disconnection between the detection results and the video screen in traditional systems, but also ensures the smoothness and real-time nature of video display through an independent video display unit, enabling users to intuitively observe the real-time status of the fire passage and the detection results of the system. At the same time, this structural design also provides good scalability for the system, facilitating the addition of new visual effects or interaction functions in the future, reflecting the innovation and practical value of this application in the design of intelligent monitoring systems.

[0099] In actual application scenarios, the synchronization problem between the front-end video stream and target rendering display has always been a technical difficulty that needs to be solved urgently. In traditional solutions, due to the differences in the processing mechanisms of target detection inference and video stream transmission, the "floating box" phenomenon where the target box is not aligned with the target object in the video stream screen often occurs. Especially when the inference frame rate of target detection cannot reach the video stream frame rate, it will cause an asynchronous problem where the position of the target box remains static while the target object in the video screen continues to move, seriously affecting the usability and user experience of the monitoring system. To address this problem, existing technologies usually adopt the method of decoding the video stream on the edge side, directly rendering the target box on the decoded image, and then re-encoding it into a video stream. Although this solution solves the synchronization problem in principle, there are many limitations in actual applications: First, edge devices generally have limited CPU performance and are difficult to support the decoding, rendering, and encoding processing of multiple video streams; second, the re-encoding process significantly increases the performance overhead of the system and affects the real-time nature of the system; finally, in the scenario where multiple users access simultaneously, since the video stream and the target box have been encoded together, it is impossible to dynamically adjust the rendering method of the target box according to the personalized needs of different users, greatly reducing the flexibility of the system.

[0100] To fundamentally solve these problems, this application innovatively proposes a video stream and target box synchronous rendering solution based on timestamp matching. In this solution, the video display unit first extracts a list of target object information from the real-time video stream. Through the timestamp query and matching mechanism, it ensures that for each frame of the video image, the corresponding target object information can be accurately found. This timestamp-based matching mechanism avoids the forced synchronization processing in the traditional solution, ensuring both rendering accuracy and reducing system overhead. More importantly, to solve the problem that the target detection frame rate is lower than the video stream frame rate, the system innovatively introduces a data caching and frame rate compensation mechanism: by caching the target information list data and real-time video stream data packets within a preset time threshold range, and combining the linear interpolation algorithm to supplement the frame rate of the target information list data, a smooth transition of the target box position is achieved. This solution not only effectively solves the synchronization problem between the target box and the video image, but also reduces the requirements for the performance of the edge device through algorithm optimization. At the same time, since the rendering process of the target box is relatively independent of the transmission process of the video stream, the system can flexibly support personalized display requirements in multi-user scenarios, such as displaying different styles of target boxes according to the permission levels or preferences of different users, greatly enhancing the practicality and interaction experience of the system. The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An interactive system for fire passage occupancy monitoring, characterized in that: include: An image acquisition device for acquiring a monitoring image of a target fire passage; A display device for interacting with a user; A server connected to the image acquisition device and the display device by signals respectively; the server is configured to: determine whether the fire passage is currently occupied based on the monitoring image of the target fire passage acquired by the image acquisition device, and if so, warn the user through the display device; The server includes a label acquisition module, a built-in AI engine, an open target detection model and an alarm module; The label acquisition module is configured to be signal-connected to the multimodal large model and the image acquisition device respectively, and is used to periodically acquire at least one frame of monitoring image of the target fire passage, input the multimodal large model for preliminary target recognition, and use a preset prompt word template to guide the multimodal large model to output a first set of labels to be detected of the monitoring image; use the first set of labels to be detected to update a second set of labels to be detected with a fixed length in chronological order, and when the second set of labels to be detected changes, notify the built-in AI engine to start; the first set of labels to be detected includes Chinese and English labels; The built-in AI engine is connected to the image acquisition device and the open target detection model signal, and is used to acquire the video stream in the current period and decode to obtain several frames of video stream images, and send the video stream images and capability parameters to the open target detection model; The open target detection model is connected to the alarm module signal, and is used to use the capability parameter as an input parameter to complete the target object detection of the video stream image and output a target object information list; The alarm module is connected to the display device signal, and is used to statically track the target object in the target fire passage interest area according to the target object information list, and output the tracked target object information list; When the tracking target object in the tracking target object information list meets the preset alarm condition, the fire passage occupation alarm event is reported to the user through the display device.

2. The interactive system for fire passage occupancy monitoring according to claim 1, characterized in that: The alarm module includes a target tracking unit and an event reporting unit; The target tracking unit is used to statically track the target object in the target fire passage interest area according to the target object information list, and output the tracked target object information list; The event reporting unit is connected to the target tracking unit and the display device by signal respectively, and is used to determine whether the tracking target object in the tracking target object information list meets the preset alarm condition, and if so, reports the alarm information to the display device.

3. The interactive system for fire passage occupancy monitoring according to claim 2, characterized in that: The alarm module includes a video fusion unit and a video display unit; The video fusion unit is respectively connected to the built-in AI engine and the open target detection model signal, and is used to obtain the video stream image and the target object information list, and output the real-time video stream and the target object identification frame to the video display unit; The video display unit is connected to the display device signal, and is used to obtain the real-time video stream and the target object recognition frame, and send them to the display device after rendering for display to the user.

4. The interactive system for fire passage occupancy monitoring according to claim 1, characterized in that: The capability parameters include: camera ID, video stream address, video stream format, area of ​​interest, AI capability name, tracking time threshold, and interval alarm period.

5. The interactive system for fire passage occupancy monitoring according to claim 1, characterized in that: The target object information list includes: target position, target category, target confidence, and image timestamp; the tracking target object information list includes: tracking ID, tracking status, target duration, target position, target category, target threshold, and image timestamp.

6. The interactive system for fire passage occupancy monitoring according to claim 2, characterized in that: The method for the target tracking unit to perform static tracking includes: Filter the target objects in the target object information list that do not belong to the target fire channel interest area; Merge target objects whose intersection-and-union ratio is greater than a preset intersection-and-union threshold among the remaining target objects; If the remaining target objects do not appear at the target location within the preset time, it is determined that the target object is lost.

7. The interactive system for fire passage occupancy monitoring according to claim 5, characterized in that: The alarm conditions include: The duration of the tracking target object in the tracking target object information list at the target location is greater than the preset time; The number of reports of the tracking target objects in the tracking target object information list within the interval alarm period shall not exceed 1 time.

8. The interactive system for fire passage occupancy monitoring according to claim 7, characterized in that: If the interval alarm period is 0, the tracking target object in the tracking target object information list is reported only once in the whole monitoring period.

9. The interactive system for fire passage occupancy monitoring according to claim 3, characterized in that: The method for aligning the target object recognition frame with the target object in the real-time video stream when the video display unit completes rendering includes: Obtain a target object information list in a real-time video stream, query the image timestamp, match and obtain the target object information list and the real-time video stream data packet with the same timestamp, render and send to a display device for display to the user.

10. The interactive system for fire passage occupancy monitoring according to claim 9, characterized in that: The rendering method includes: if the frame rate of the target object information list with the same timestamp is lower than the frame rate of the real-time video stream, then caching the target information list data and the real-time video stream data packet within a preset time threshold range, and using a linear interpolation method to supplement the frame rate of the cached target information list data.

Citation Information

Patent Citations

  • Behavior recognition method and device based on multi-modal large model and electronic equipment

    CN118314624A

  • Target tracking method based on fusion of classic recognition model and multi-modal large model

    CN119004388A