Unattended system
By combining a cloud-based image analysis and processing center with an AI language model, the problems of lagging fire alarm recognition and poor device linkage compatibility in traditional fire alarm monitoring systems have been solved. This enables automatic capture, accurate judgment, and multi-terminal linkage of fire alarm behavior, improving the efficiency and accuracy of fire alarm monitoring systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN STARCAM TECH
- Filing Date
- 2026-03-03
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional fire alarm monitoring systems are susceptible to environmental interference leading to false alarms, have delayed fire alarm identification, consume high bandwidth for video transmission, have poor device linkage compatibility, lack fire information, and have low emergency response efficiency.
By employing a cloud-based image analysis and processing center combined with an AI language model, the system achieves automated analysis of fire alarm features and semantic alarms. It also enables precise audible and visual alarms through linkage with alarm devices, solving problems such as high video transmission bandwidth consumption, delayed fire alarm recognition, and poor device linkage compatibility. This allows for the automatic capture, accurate judgment, and multi-terminal linkage response of fire alarm behaviors.
It enables automated analysis and accurate judgment of fire alarm behavior, reduces video transmission bandwidth usage, improves fire alarm linkage response efficiency, provides accurate key fire information, and reduces equipment deployment and maintenance costs.
Smart Images

Figure CN122024397A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of monitoring technology, specifically to unattended systems. Background Technology
[0002] Traditional fire alarm monitoring systems have many technical drawbacks: sensors are easily affected by environmental interference, resulting in false alarms. For example, smoke and steam can trigger smoke sensors, while dust and high temperatures can affect the accuracy of temperature sensors, leading to frequent invalid alarms. Video monitoring footage relies on real-time manual analysis, making it difficult for monitoring personnel to continuously monitor multiple video streams. This can easily lead to the omission of key fire alarm features such as open flames, smoke, and sparks, resulting in delayed fire alarm identification and missed opportunities for optimal response. The real-time video streams collected by video surveillance equipment are raw, unstructured data. Uploading them directly to the cloud will consume a lot of network bandwidth, resulting in low transmission efficiency. Furthermore, the massive amount of video data lacks targeted filtering, leading to high computing costs for subsequent data analysis. The determination of fire alarm types, classification of levels, and generation of handling instructions rely entirely on human experience, making it impossible to achieve automated analysis and accurate judgment of fire alarm behavior. Furthermore, the alarm information consists only of simple sound and light signals, lacking key information such as the location, type, and level of the fire, resulting in low emergency response efficiency. The protocols of on-site fire alarm equipment and emergency equipment at the remote fire management terminal are not consistent, and the problem of equipment heterogeneity is prominent. It is difficult to achieve cross-regional and cross-type fire alarm linkage and alarm, and it is impossible to form a closed loop of rapid on-site reminder and remote emergency response. Summary of the Invention
[0003] This application proposes an unattended system that addresses the technical shortcomings of traditional fire alarm monitoring by employing a cloud-based image analysis and processing center to achieve centralized and automated analysis of fire alarm features. Combined with the semantic analysis and reasoning capabilities of an AI language model, it converts fire alarm image features into natural language alarm information containing key information. Through linkage with alarm devices, it achieves accurate audible and visual alarms at the scene and remotely. This solves problems such as high video transmission bandwidth consumption, delayed fire alarm recognition, ambiguous alarm information, and poor device linkage compatibility, enabling automatic capture, accurate judgment, semantic alarm, and multi-terminal linkage response of fire alarm behavior.
[0004] This application provides an unattended system, including: an AI language model, video surveillance equipment, a cloud-based image analysis and processing center, and a linkage alarm device; wherein, the video surveillance equipment is connected to the cloud-based image analysis and processing center via network communication, the cloud-based image analysis and processing center is connected to the AI language model via network communication, the AI language model is connected to the linkage alarm device via network communication, the video surveillance equipment is configured with an audible and visual alarm component, and the linkage alarm device includes multiple protocol conversion middleware; The video surveillance equipment is used to collect real-time video streams of the current monitored area, and upload the real-time video streams to the cloud image analysis and processing center after lightweight encoding. The AI language big model is used to analyze the target image features extracted by the cloud image analysis and processing center, determine the fire alarm behavior of the illegal fire alarm event, and generate natural language alarm information; wherein, the fire alarm behavior is determined by the cloud image analysis and processing center based on the fire alarm feature capture of real-time video stream and the data analysis of the cloud computing power cluster.
[0005] Optionally, in some embodiments of this application, the video surveillance device is used to determine a first region on the current frame image of the real-time video stream, so as to determine a first target identifier of the first focusing target on the current frame image; wherein, the first target identifier is a positioning identifier of the second identifier information of any target fire alarm supervision behavior and fire alarm characteristics in the preset fire alarm supervision behavior, and the second identifier information is used to determine the target frame image of the fire alarm behavior of the fire alarm target on the real-time video stream of the current monitoring area, and the high-definition panoramic camera encodes the first region frame data with the first target identifier and uploads it separately to the cloud image analysis and processing center.
[0006] Optionally, in some embodiments of this application, the first region is configured with a first bounding box and a second bounding box; wherein, the first bounding box is a display box of the fire alarm behavior on the current frame image, the second bounding box is a display box of the fire alarm source in the fire alarm behavior, and the second bounding box is configured with a millisecond-level timestamp; The target frame image is time-stamped, and the high-definition panoramic camera uploads the time-stamped target frame image to the cloud image analysis and processing center after lightweight compression. The cloud image analysis and processing center fuses the target frame images sequentially by timestamp to generate structured real-time video stream data.
[0007] Optionally, in some embodiments of this application, the cloud image analysis and processing center is configured with a cloud distributed storage module, which stores data features of fire alarm behavior. The data features of fire alarm behavior include behavioral feature identifiers in chronological order. The initial identifier information of the behavioral feature identifier is the first frame compressed image in the real-time video stream when the fire alarm feature is the target fire alarm. The sequentially adjacent images of the first frame image are the target frame compressed images corresponding to the data features of the fire alarm behavior recognized by the cloud image analysis and processing center, which contain the target fire alarm and the fire alarm feature.
[0008] Optionally, in some embodiments of this application, the cloud image analysis and processing center is deployed with a cloud-distributed visual feature extraction model. This model runs on a cloud computing cluster and is used to determine the fire alarm target image features corresponding to the target frame image in structured real-time video stream data, and to calculate the first confidence value of the fire alarm features in the target frame image belonging to the archived fire type and the second confidence value of the fire alarm behavior features in the target frame image belonging to the preset fire level under the corresponding fire type. The first confidence level and the second confidence level are fused to determine the fused confidence level value. When the fused confidence level value is higher than the first threshold, the target frame image is used as a fire alarm judgment image with fire alarm features and pushed to the AI language big model.
[0009] Optionally, in some embodiments of this application, the cloud-based image analysis and processing center is configured with a second threshold and a third threshold; wherein, when the first confidence value is lower than the second threshold, the cloud-based image analysis and processing center sends an instruction to the linkage alarm device, which then performs a fire alarm identification anomaly alarm; when the second confidence value is lower than the third threshold and the first confidence value is higher than the second threshold, the cloud-based image analysis and processing center sends an instruction to the linkage alarm device, which then performs a fire level matching anomaly alarm.
[0010] Optionally, in some embodiments of this application, the AI language big model includes a multimodal input interface and a big language processing model. The multimodal input interface includes a video analysis input interface and a fire system data input interface. The video analysis input interface is used to receive a first representation text pushed by a cloud-based image analysis and processing center, based on the monitoring area level and fire duty label. The first representation text is the fire alarm handling task text corresponding to the fire alarm feature and the first fire handling type under the monitoring area level. The fire system data input interface is used to receive a second representation text corresponding to the fire alarm handling task text. The second representation text is the fire alarm behavior handling text under the current fire alarm handling task for the fire alarm feature and the corresponding fire alarm type. The large language processing model is used to generate natural language alarm information with causal relationship based on the first and second representation texts, and to execute the fire alarm under the current fire alarm handling task according to the fire type.
[0011] Optionally, in some embodiments of this application, the natural language alarm information includes a first alarm information that integrates cloud-based spatiotemporal positioning data, based on time, fire location, and monitoring area, and a second alarm information that integrates real-time data from the cloud-based fire protection system, based on fire status, risk level, and fire alarm operation and maintenance.
[0012] Optionally, in some embodiments of this application, the large language processing model further determines the first fire alarm equipment information of the monitoring area corresponding to the current fire alarm handling task based on the current fire alarm handling task of the fire type and combined with the spatial positioning data of the cloud image analysis and processing center; wherein, the first fire alarm equipment information consists of video surveillance equipment information with sound and light alarm function and second fire alarm equipment information in the current monitoring area whose distance from the fire source does not exceed a preset maximum alarm distance threshold.
[0013] Optionally, in some embodiments of this application, the linkage alarm device is further configured with a field alarm control unit based on a first protocol stack and a remote linkage alarm control unit based on a second protocol stack, and both the first protocol stack and the second protocol stack support cloud communication protocol adaptation with the cloud image analysis and processing center and the AI language big model; wherein, the first protocol stack is used to receive a first protocol configuration instruction issued by the AI language big model based on protocol conversion configuration, which has the permission to call the device in the current monitoring area, and the first protocol configuration instruction is a control instruction for the audible and visual alarm signals of the linkage alarm device and video surveillance device in the current monitoring area; the second protocol stack is used to receive a second protocol configuration instruction issued by the AI language big model based on protocol conversion configuration, which has a remote request, and the second protocol configuration instruction is a control instruction for the audible and visual alarm signals of the mobile alarm device and emergency fire response device at the remote fire management terminal.
[0014] This application provides an unattended system, comprising: an AI language model, video surveillance equipment, a cloud-based image analysis and processing center, and a linked alarm device; wherein, the video surveillance equipment is connected to the cloud-based image analysis and processing center via network communication, the cloud-based image analysis and processing center is connected to the AI language model via network communication, the AI language model is connected to the linked alarm device via network communication, the video surveillance equipment is configured with an audible and visual alarm component, and the linked alarm device includes multiple protocol conversion middleware; the video surveillance equipment is used to acquire real-time video streams of the current monitored area, and uploads the real-time video streams to the cloud-based image analysis and processing center after lightweight encoding; the AI language model is used to analyze the target image features extracted by the cloud-based image analysis and processing center, determine the fire alarm behavior of an illegal fire alarm event, and generate natural language alarm information; wherein, the fire alarm behavior is determined by the cloud-based image analysis and processing center based on real-time video... Based on the capture of fire alarm features and data analysis of cloud computing clusters, the unattended system provided in this application first performs lightweight encoding on the real-time video stream, effectively reducing the size of video data, reducing network bandwidth consumption, and improving cloud transmission efficiency. The cloud image analysis and processing center relies on the distributed computing capabilities of the cloud computing cluster to achieve simultaneous capture and analysis of fire alarm features in multiple areas, replacing traditional edge computing devices, reducing on-site equipment deployment and maintenance costs, and the centralized cloud processing enables unified management of fire alarm feature data and iterative optimization of models. The AI language big model transforms the fire alarm image features extracted from the cloud into structured, semantic natural language alarm information, breaking through the information limitations of traditional simple sound and light alarms and providing accurate key fire information for fire response. The linkage alarm devices solve the problem of device heterogeneity through protocol conversion middleware, realizing unified scheduling of alarm devices of different brands and types, and improving the efficiency of fire alarm linkage response. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the structure of the monitoring system provided in the embodiments of this application. Detailed Implementation
[0017] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0018] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0019] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0020] The following describes in detail the embodiments involved in this application. It should be noted that the order of description of the embodiments in this application is not intended to limit the priority of the embodiments.
[0021] Figure 1 This is a schematic diagram of the architecture of an unattended system provided in this application embodiment. The system includes: an AI language big data model, video surveillance equipment, a cloud image analysis and processing center, and linkage alarm equipment. Each device is interconnected in the cloud through a wired / wireless communication network. The video surveillance equipment is deployed in key fire hazard points (such as power distribution rooms, warehouses, kitchens, etc.) in monitoring areas such as industrial plants, commercial complexes, and parks. It adopts high-definition panoramic cameras and is equipped with sound and light alarm components, supporting lightweight video encoding and network uploading. The cloud-based image analysis and processing center is deployed on a public / private cloud platform, configured with a cloud computing cluster (distributed GPU / CPU computing power), a cloud-based distributed storage module, and a cloud-based distributed visual feature extraction model to achieve centralized analysis, storage, and fire alarm feature extraction of video stream data; the AI language big model is deployed on the same cloud platform as the cloud-based image analysis and processing center, supporting multimodal data input and natural language generation to achieve semantic conversion of fire alarm information; The linkage alarm equipment includes on-site alarm devices (such as audible and visual alarms and fire broadcasts), remote fire management terminal devices (such as mobile terminals, fire emergency command platforms, and drones), and multiple protocol conversion middleware to realize on-site and remote fire alarm linkage.
[0022] The core workflow of the system is as follows: video surveillance equipment acquires real-time video streams and performs H.265 / HEVC lightweight encoding, then uploads them to the cloud-based image analysis and processing center via the network; the cloud-based image analysis and processing center uses a computing cluster to capture fire alarm features, extract target image features, complete dual confidence calculations for fire type and level, and pushes images that meet the threshold requirements for fire alarm determination to the AI language model; the AI language model integrates cloud-based image analysis data and fire protection system business data to generate natural language alarm information and selects the optimal combination of alarm devices; the linked alarm devices adapt to cloud commands through protocol conversion middleware, driving on-site and remote alarm devices to generate audible and visual alarm signals, achieving accurate fire alarm and coordinated response.
[0023] The video surveillance equipment uses high-definition panoramic cameras and has a built-in fire alarm feature recognition algorithm, which can identify fire alarm features such as open flames, smoke, and sparks in the video stream in real time. In the current frame of the real-time video stream captured by the camera, the identified fire alarm features are located at the pixel level to determine the first region containing complete fire alarm features. This region is the area with the least fire feature and avoids the inclusion of invalid image areas. Based on a pre-set fire alarm monitoring behavior database (such as open flame combustion, smoke diffusion, sparks, etc.) and a fire alarm feature database (such as the color and shape of flames, the concentration and texture of smoke, etc.), a second identification information is generated. This information is a matching identifier between the fire alarm feature and the pre-set fire alarm monitoring behavior. Based on the second identification information, a first target identifier is generated. This identifier is a positioning identifier that includes the location and type of fire alarm features (such as coordinates + fire feature code). The camera encodes the first area frame data with the first target identifier separately (separated from the whole frame video stream) and uploads it separately to the cloud image analysis and processing center through the network to achieve accurate transmission of fire alarm feature data.
[0024] Within the first area defined by the video surveillance equipment, a dual bounding box is dynamically marked using a target detection algorithm: the first bounding box is the fire alarm behavior box, marking the development behavior of the fire (such as the spread range of flames and the area of smoke dispersion), which is dynamically adjusted in real time as the fire develops; the second bounding box is the fire alarm source box, marking the source of the fire (such as the equipment or items that are on fire), which is a fixed location box. Configure Unix millisecond-level timestamps for the second bounding box, and time-stamp each frame of target image containing fire alarm features. The timestamps are synchronized with the camera's system time to ensure time consistency across devices. The video surveillance equipment performs lightweight compression (using a lossless compression algorithm to ensure feature integrity) on the target frame images marked with double bounding boxes and millisecond-level timestamps, and uploads them to the cloud image analysis and processing center. The cloud image analysis and processing center performs time-series sorting and fusion of the received target frame images based on the timestamps to generate structured real-time video stream data containing the fire development process, enabling continuous analysis of fire alarm behavior.
[0025] The cloud-based distributed storage module uses a distributed file system to classify and store data characteristics of fire alarm behavior according to fire type (such as open flame, smoke, electrical fire). Each fire type corresponds to a unique behavior feature identifier, which contains time-series fire alarm development feature data. When the cloud receives the target frame image uploaded by the video surveillance device, it first identifies the fire alarm features in the image, determines the target fire alarm type, and uses the first frame compressed image of the fire alarm as the initial identification information of the behavior feature identifier, and stores it in the feature library of the corresponding fire type. Based on the first frame image, subsequent uploaded frame images are filtered through a temporal matching algorithm. Sequentially adjacent images that contain the same fire alarm features and conform to the development pattern of fire alarm behavior are used as target frame compressed images and associated with the initial identification information to form a complete fire alarm behavior temporal feature chain. The cloud-based distributed storage module only stores compressed images of target frames associated with behavioral feature identifiers. Unmatched frame images will be automatically cleaned up, enabling accurate filtering and efficient storage of fire alarm data.
[0026] The cloud-based distributed visual feature extraction model is trained based on deep learning algorithms and runs on a cloud computing cluster, supporting parallel feature extraction of multiple frames of images. The model extracts fire target image features (such as flame texture, smoke concentration, and fire spread rate) from the structured real-time video stream data generated in the cloud. The dual-confidence calculation process is initiated as follows: First confidence calculation: The extracted fire alarm features are matched with the fire type features archived in the cloud to calculate the matching similarity, resulting in a first confidence value (ranging from 0 to 1, with higher values indicating a higher degree of fire type matching). Second confidence calculation: Based on the determined fire type, the fire alarm behavior features are matched with preset fire level features for that type (e.g., Level 1 fire: small flames, no spread; Level 2 fire: flame spread, moderate smoke concentration; Level 3 fire: large-area flames, dense smoke), to calculate the matching similarity, resulting in a second confidence value (ranging from 0 to 1, with higher values indicating a higher degree of fire level matching). A weighted fusion algorithm is used to fuse the first confidence level and the second confidence level to obtain a fused confidence level value. The preset first threshold is 0.8. When the fused confidence level value is higher than 0.8, the image is determined to be a fire alarm image with fire alarm characteristics and is pushed to the AI language big model. If it is lower than 0.8, the monitoring of subsequent video streams continues.
[0027] The cloud-based image analysis and processing center is configured with a second threshold and a third threshold, where the second threshold is preset to 0.7 (fire type matching threshold) and the third threshold is preset to 0.75 (fire level matching threshold). When the first confidence value is lower than the second threshold (<0.7), it indicates that the fire alarm features cannot be effectively matched with the archived fire types, and there is an abnormality in fire alarm identification (such as suspected fire or blurred features). The cloud image analysis and processing center immediately sends an instruction to the linkage alarm device, which then executes the fire alarm identification abnormality alarm. The on-site alarm device issues an audible and visual prompt of "Fire feature identification abnormal, please confirm manually". When the first confidence value is higher than the second threshold (≥0.7), but the second confidence value is lower than the third threshold (<0.75), it indicates that the fire type can be accurately identified, but the fire level cannot be effectively matched. The cloud image analysis and processing center sends a command to the linkage alarm device, which then executes an alarm for abnormal fire level matching. The on-site alarm device issues an audible and visual prompt: "Fire type has been identified, but the level cannot be matched. Please make a manual judgment." When the first confidence value is ≥0.7 and the second confidence value is ≥0.75, it indicates that the fire type and level are effectively matched. The cloud will push the relevant data to the AI language big model and enter the natural language alarm information generation stage.
[0028] The AI language large model is configured with multimodal input interfaces, which are divided into video analysis input interface and fire protection system data input interface. Both interfaces support real-time data connection with cloud systems. The video analysis input interface receives data analysis results pushed by the cloud image analysis and processing center. Based on the monitoring area level (such as level 1 protection zone, level 2 protection zone) and fire duty label (such as power distribution room fire protection, warehouse fire protection), it generates the first characterization text. This text is a structured business text, for example: "Level 1 protection zone - power distribution room, open flame characteristics detected, corresponding fire handling type is initial handling of electrical fire". The fire protection system data input interface connects with the fire emergency response system and receives the second characterization text corresponding to the first characterization text. This text is the standardized handling procedure text for this type of fire, for example: "In the initial handling of electrical fires, immediately cut off the power supply, use a dry powder fire extinguisher to extinguish the fire, and do not use water to extinguish the fire." The AI language big model performs causal correlation analysis on the first and second representation texts, and combines the basic information of the fire to generate natural language alarm information with causal correlation. For example: "An open flame was detected in the power distribution room of the first protection zone. It is determined to be an initial electrical fire. The power supply to the power distribution room should be cut off immediately. Use a dry powder fire extinguisher to put it out. Do not use water to put it out."
[0029] The natural language alarm information generated by the AI language model is divided into first alarm information and second alarm information. The information from the two dimensions is integrated with cloud spatiotemporal positioning data and real-time data from the fire protection system to form a complete alarm information chain. First alarm information (basic spatiotemporal dimension): Integrating spatial positioning data (camera coordinates, fire pixel location) and timestamp data from the cloud image analysis and processing center, basic information including time, fire location, and monitoring area is generated. For example: "On [Date] at [Time], a fire broke out in the power distribution room of XX Park (camera number: SY001, coordinates: X / Y)". Second alarm information (business handling dimension): Integrate real-time data from the cloud-based fire protection system (fire status, risk level, operation and maintenance database) to generate business information that includes fire status, risk level, and fire alarm operation and maintenance. For example: "The fire status is the initial stage of open flame with no spread. The risk level is level two. It is recommended to immediately cut off the power supply to the power distribution room, use dry powder fire extinguishers to put out the fire, and evacuate the personnel on site to a safe area quickly." The AI language big data model seamlessly integrates the first and second alarm information to generate complete natural language alarm information containing basic and business information, realizing integrated information output for fire "location-judgment-handling".
[0030] The AI language big model's big language processing model first retrieves the fire equipment deployment information of the monitoring area corresponding to the current fire alarm handling task based on the fire type from the fire protection system. By combining spatial positioning data from the cloud-based image analysis and processing center, the precise coordinates of the fire alarm source are determined, and video surveillance equipment with audible and visual alarm functions (i.e., high-definition panoramic cameras with audible and visual alarm components deployed near the fire alarm source) are selected as basic alarm equipment. The maximum alarm distance threshold is preset to 50 meters. Based on the coordinates of the fire alarm source, a second fire alarm device (such as an on-site audible and visual alarm or fire broadcast) within the monitoring area that is no more than 50 meters away from the fire alarm source is selected as an auxiliary alarm device. The combination of information from basic alarm devices and auxiliary alarm devices constitutes the first fire alarm device information. The AI language big data model sends this information, along with natural language alarm information, to the linked alarm devices to achieve precise scheduling of alarm devices.
[0031] The linkage alarm device is equipped with a field alarm control unit and a remote linkage alarm control unit, which are respectively equipped with a first protocol stack and a second protocol stack. Both protocol stacks have built-in cloud communication protocol adaptation modules, which support protocol docking with cloud image analysis and processing centers and AI language large models. The first protocol stack uses field communication protocols such as Modbus / CAN bus to receive the first protocol configuration command issued by the AI language big model. This command is the sound and light alarm control command of the field alarm device. After being adapted by the protocol conversion middleware, it drives the sound and light alarm components of the video surveillance equipment and the second fire alarm device near the fire source to issue sound and light alarm signals containing natural language (such as playing a complete natural language alarm message and the light flashing red), so as to realize the rapid and accurate alarm on site. The second protocol stack adopts cloud communication protocols such as MQTT / HTTP / HTTPS to receive the second protocol configuration instructions issued by the AI language big model. These instructions are emergency equipment linkage instructions from the remote fire management terminal. After being adapted by the protocol conversion middleware, natural language alarm information is pushed to devices such as the fire emergency command platform, firefighter mobile terminals, and fire drones, and the sound and light alarms of the remote devices are triggered to achieve rapid emergency response from the remote fire management terminal. The middleware converts the unified JSON format commands issued by the AI language model into private protocol commands for different brands and types of devices, enabling seamless device integration. At the same time, the middleware can dynamically sense the device status (online / offline) and automatically select the optimal communication link to ensure stable and fast transmission of alarm commands.
[0032] The above provides a detailed description of a monitoring system provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An unattended system, characterized in that, include: The system comprises an AI language model, video surveillance equipment, a cloud-based image analysis and processing center, and a linked alarm device. The video surveillance equipment is connected to the cloud-based image analysis and processing center via a network, the cloud-based image analysis and processing center is connected to the AI language model via a network, the AI language model is connected to the linked alarm device via a network, the video surveillance equipment is equipped with an audible and visual alarm component, and the linked alarm device includes multiple protocol conversion middleware. The video surveillance equipment is used to collect real-time video streams of the current monitored area, and upload the real-time video streams to the cloud image analysis and processing center after lightweight encoding. The AI language big model is used to analyze the target image features extracted by the cloud image analysis and processing center, determine the fire alarm behavior of the illegal fire alarm event, and generate natural language alarm information; wherein, the fire alarm behavior is determined by the cloud image analysis and processing center based on the fire alarm feature capture of real-time video stream and the data analysis of the cloud computing power cluster.
2. The unattended system according to claim 1, characterized in that, The video surveillance device is used to determine a first region on the current frame image of the real-time video stream, so as to determine a first target identifier of a first focusing target on the current frame image; wherein, the first target identifier is a positioning identifier of the second identifier information of any target fire alarm supervision behavior and fire alarm characteristics in the preset fire alarm supervision behavior, and the second identifier information is used to determine the target frame image of the fire alarm behavior of the fire alarm target on the real-time video stream of the current monitoring area, and the high-definition panoramic camera encodes the first region frame data with the first target identifier and uploads it separately to the cloud image analysis and processing center.
3. The unattended system according to claim 2, characterized in that, The first region is configured with a first bounding box and a second bounding box; wherein, the first bounding box is the display box of the fire alarm behavior on the current frame image, the second bounding box is the display box of the fire alarm source in the fire alarm behavior, and the second bounding box is configured with a millisecond-level timestamp; The target frame image is time-stamped, and the high-definition panoramic camera uploads the time-stamped target frame image to the cloud image analysis and processing center after lightweight compression. The cloud image analysis and processing center fuses the target frame images sequentially by timestamp to generate structured real-time video stream data.
4. The unattended system according to claim 1, characterized in that, The cloud-based image analysis and processing center is equipped with a cloud-based distributed storage module, which stores data features of fire alarm behavior. These data features include behavioral feature identifiers arranged in chronological order. The initial identifier information of the behavioral feature identifier is the first compressed image of the real-time video stream when the fire alarm feature is the target fire. The sequentially adjacent images of the first frame are the target frame compressed images corresponding to the data features of the fire alarm behavior recognized by the cloud-based image analysis and processing center, which contain the target fire and the fire alarm feature.
5. The unattended system according to claim 4, characterized in that, The cloud-based image analysis and processing center is equipped with a cloud-based distributed visual feature extraction model. This model runs on a cloud computing cluster and is used to determine the fire alarm target image features corresponding to the target frame image in structured real-time video stream data. It also calculates the first confidence value that the fire alarm features in the target frame image belong to the archived fire type and the second confidence value that the fire alarm behavior features in the target frame image belong to the preset fire level under the corresponding fire type. The first confidence level and the second confidence level are fused to determine the fused confidence level value. When the fused confidence level value is higher than the first threshold, the target frame image is used as a fire alarm judgment image with fire alarm features and pushed to the AI language big model.
6. The unattended system according to claim 1, characterized in that, The cloud-based image analysis and processing center is configured with a second threshold and a third threshold. When the first confidence value is lower than the second threshold, the cloud-based image analysis and processing center sends a command to the linkage alarm device, which then performs an alarm for fire alarm identification anomaly. When the second confidence value is lower than the third threshold and the first confidence value is higher than the second threshold, the cloud-based image analysis and processing center sends a command to the linkage alarm device, which then performs an alarm for fire level matching anomaly.
7. The unattended system according to claim 1, characterized in that, The AI language big model includes a multimodal input interface and a big language processing model. The multimodal input interface includes a video analysis input interface and a fire system data input interface. The video analysis input interface is used to receive a first representation text based on the monitoring area level and fire duty label pushed by the cloud image analysis and processing center. The first representation text is the fire alarm handling task text corresponding to the fire alarm feature and the first fire handling type under the monitoring area level. The fire system data input interface is used to receive a second representation text corresponding to the fire alarm handling task text. The second representation text is the fire alarm behavior handling text under the current fire alarm handling task for the fire alarm feature and the fire alarm type. The large language processing model is used to generate natural language alarm information with causal relationship based on the first and second representation texts, and to execute the fire alarm under the current fire alarm handling task according to the fire type.
8. The unattended system according to claim 7, characterized in that, The natural language alarm information includes a first alarm information that integrates cloud-based spatiotemporal positioning data, based on time, fire location, and monitoring area, and a second alarm information that integrates real-time data from the cloud-based fire protection system, based on fire status, risk level, and fire alarm operation and maintenance.
9. The unattended system according to claim 7, characterized in that, The large language processing model also determines the information of the first fire alarm device in the monitoring area corresponding to the current fire alarm handling task based on the current fire type and the spatial positioning data of the cloud image analysis and processing center. The information of the first fire alarm device consists of the information of video surveillance equipment with sound and light alarm functions and the information of the second fire alarm device in the current monitoring area whose distance from the fire source does not exceed the preset maximum alarm distance threshold.
10. The unattended system according to claim 1, characterized in that, The linkage alarm device is also equipped with a field alarm control unit based on a first protocol stack and a remote linkage alarm control unit based on a second protocol stack. Both the first and second protocol stacks support cloud communication protocol adaptation with the cloud image analysis and processing center and the AI language big model. The first protocol stack is used to receive a first protocol configuration instruction issued by the AI language big model based on protocol conversion configuration, which contains the device permission call of the current monitoring area. The first protocol configuration instruction is the control instruction for the audible and visual alarm signals of the linkage alarm device and video surveillance device in the current monitoring area. The second protocol stack is used to receive a second protocol configuration instruction issued by the AI language big model based on protocol conversion configuration, which contains the remote request. The second protocol configuration instruction is the control instruction for the audible and visual alarm signals of the mobile alarm device and emergency fire response device at the remote fire management terminal.