Large Model-Based Fire Lane Occupancy Warning Level Analysis System and Its Method
By applying a large-model-based early warning system in the fire passage, real-time monitoring and analysis of the state of the fire passage, the problems of misjudgment of traditional detection methods and unscientific standards are solved, and the safety and emergency response speed of the fire passage are improved.
Patent Information
- Application Number
- CN202510224920.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional fire passage occupation detection methods are easily affected by factors such as light and angle, and cannot capture the position information of objects changing over time. The warning level standards lack flexibility and scientific basis.
The fire passage occupation warning level analysis system is adopted based on the large model, and the fire passage status is monitored in real time through the camera, and the visual big model is used to detect the fire passage status, object position and changes, and the warning level analysis is carried out in combination with the large language model.
It improves the speed and reliability of fire emergency response, reduces safety hazards caused by blockage of fire passages, and ensures that fire passages are always open.
Smart Images

Figure CN119723469B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fire passage occupancy warning, and more specifically, to a large model-based fire passage occupancy warning level analysis system and method thereof. Background Art
[0002] With the acceleration of the urbanization process and the increase in building density, fire safety has become increasingly important. As an important path for firefighters to quickly reach the fire scene in case of an emergency, the unobstructedness of the fire passage is the key to ensuring the safety of life and property. However, in daily management, the phenomenon of fire passages being occupied is not uncommon, which not only increases the difficulty of rescue in case of a fire but also poses a potential threat to the lives of residents.
[0003] Traditional fire passage occupancy detection usually relies on analyzing a single image captured by a camera to detect whether there is an object in the target area to determine whether there is occupancy. This method is easily affected by factors such as lighting and angle, resulting in misjudgment. In addition, it cannot capture the position information of objects changing over time, which may lead to misinterpretation of dynamic scenes. Moreover, past practices often required pre-setting multiple safety levels, such as which type of target occupancy is considered a high risk, etc. These criteria are mostly based on human experience, lacking flexibility and scientific basis. With the change of environmental conditions, the original criteria may no longer apply, thus affecting the judgment accuracy. Summary of the Invention
[0004] This application provides a large model-based fire passage occupancy warning level analysis system and method thereof, which can automatically identify and analyze the status of the fire passage, timely detect abnormalities and issue alarms, so as to ensure that the fire passage is always unobstructed.
[0005] In the first aspect, a large model-based fire passage occupancy warning level analysis method is provided, including:
[0006] Obtain the time series of key frames for monitoring the status of the fire passage collected by the camera;
[0007] Use the trained vision large model to detect each key frame in the time series of key frames for monitoring the status of the fire passage to obtain a time series of detection results, and the detection results include the status of the fire passage, the position of the object, and the change situation;
[0008] After adding a prompt word to the end of the time series of the detection results, input it into the fire lane occupancy warning level analysis engine based on a large language model to obtain the fire lane occupancy warning level analysis result. The prompt word is "Please judge whether there is occupancy in the fire lane according to the above description and give the following information: a. Whether there is occupancy; b. If there is occupancy, what is the duration; c. Give a reasonable warning level according to the occupancy duration and the type of the occupied object".
[0009] In a second aspect, a large model-based fire lane occupancy warning level analysis system is provided, including:
[0010] A fire lane status monitoring and acquisition module, configured to obtain the time series of key frames of the fire lane status monitoring collected by a camera;
[0011] A fire lane status detection module, configured to use a trained vision large model to detect each key frame in the time series of the key frames of the fire lane status monitoring to obtain the time series of detection results, where the detection results include the fire lane status, the object position, and the change situation;
[0012] A fire lane occupancy warning level analysis module, configured to add a prompt word to the end of the time series of the detection results and then input it into the fire lane occupancy warning level analysis engine based on a large language model to obtain the fire lane occupancy warning level analysis result. The prompt word is "Please judge whether there is occupancy in the fire lane according to the above description and give the following information: a. Whether there is occupancy; b. If there is occupancy, what is the duration; c. Give a reasonable warning level according to the occupancy duration and the type of the occupied object".
[0013] A large model-based fire lane occupancy warning level analysis system and method provided by the present application obtain the time series of key frames of the fire lane status monitoring by real-time monitoring and acquisition of the fire lane status through a camera, and then use a vision large model to detect these key frames of the fire lane status monitoring to identify the fire lane status, the object position, and the change situation in each key frame, so as to generate corresponding detection results. Finally, a large language model and prompt information are used to perform fire lane occupancy warning level analysis on these analyzed detection results to obtain the analysis result. Compared with the traditional single-frame object detection analysis result, this helps to improve the speed and reliability of fire emergency response, thereby reducing the safety hazards caused by fire lane blockage. Description of the Drawings
[0014] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings of the embodiments of the present application will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present application and do not limit the present application.
[0015] Figure 1 It is a schematic flowchart of the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0016] Figure 2 It is a schematic diagram of data flow of the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0017] Figure 3 It is a schematic flowchart of S3 in the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0018] Figure 4 It is a schematic flowchart of S32 in the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0019] Figure 5 It is a schematic flowchart of S321 in the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0020] Figure 6 It is a schematic flowchart of S322 in the method for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application.
[0021] Figure 7 It is a schematic block diagram of the system for analyzing the warning level of fire passage occupancy based on a large model in the embodiments of the present application. Detailed implementation manners
[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts also belong to the scope of protection of the present application.
[0023] In view of the above technical problems, in order to effectively monitor and warn of the occupancy of fire passages, an intelligent monitoring system based on artificial intelligence technology has emerged. Such systems can automatically identify and analyze the status of fire passages, detect abnormalities in a timely manner and issue alarms, so as to ensure that fire passages are always kept unobstructed.
[0024] Based on this, in the technical solution of this application, a method for analyzing the warning level of fire passage occupancy based on large models is proposed, aiming to use advanced machine learning technologies, including the combination of vision large models and large language models, to provide a more reliable, flexible and intelligent solution. Specifically, the technical concept of this application is to obtain the time series of key frames for monitoring the status of the fire passage by real-time monitoring and collection of the fire passage status through a camera, and then use the vision large model to detect these key frames for monitoring the fire passage status, so as to identify the fire passage status, object position and change situation in each key frame, thereby generating corresponding detection results. Finally, the large language model and prompt information are used to analyze the warning level of fire passage occupancy for these analyzed detection results to obtain the analysis results. Compared with the analysis results of traditional single-frame object detection, this helps to improve the speed and reliability of fire emergency response, thereby reducing the safety hazards caused by blocked fire passages.
[0025] Specifically, as Figure 1 and Figure 2 shown, the steps of the method for analyzing the warning level of fire passage occupancy based on large models include: S1, obtaining the time series of key frames for monitoring the fire passage status collected by the camera; S2, using the trained vision large model to detect each key frame for monitoring the fire passage status in the time series of key frames for monitoring the fire passage status to obtain the time series of detection results, and the detection results include the fire passage status, object position and change situation; S3, after adding a prompt word to the tail of the time series of detection results, inputting it into the analysis engine for the warning level of fire passage occupancy based on the large language model to obtain the analysis result of the warning level of fire passage occupancy, and the prompt word is "Please judge whether there is occupancy in the fire passage according to the above description, and give the following information: a. Whether there is occupancy; b. If there is occupancy, what is the duration; c. Give a reasonable warning level according to the occupancy duration and the type of occupied object".
[0026] Exemplarily, in step S1, the time series of key frames for monitoring the fire passage status collected by the camera is obtained. It should be understood that analyzing the situation of the fire passage through continuous key frames instead of single pictures can provide richer information to help the system understand the dynamic changes of objects in the fire passage, such as the movement, appearance or disappearance of objects, etc., so as to more reliably judge whether there is an occupancy situation. This method helps to reduce misjudgments caused by light changes, angle problems or other instantaneous factors, improve the accuracy of detection results, and better adapt to complex situations in different scenarios.
[0027] In one embodiment, obtaining the time series of the key frames for monitoring the status of the fire passage collected by the camera includes: installing appropriate camera devices around the fire passage, and these devices should have sufficient resolution and frame rate to ensure clear and coherent pictures are captured. The positions of the cameras should be carefully selected to ensure coverage of all important areas, especially those prone to being occupied. Once the cameras start working, they continuously record video streams.
[0028] Exemplarily, in the step S2, the trained large vision model is used to detect each key frame in the time series of the key frames for monitoring the status of the fire passage to obtain the time series of detection results, and the detection results include the status of the fire passage, the position of the object, and the change situation. It should be understood that traditional methods rely on single-frame image analysis, are easily affected by factors such as lighting and angle, and cannot capture the position information of objects changing over time, which may lead to misjudgments. By processing the key frames of the time series through the large vision model, these problems can be overcome, providing more reliable and detailed results.
[0029] To ensure that the large vision model can effectively process the fire passage scenario, first, a high-quality dataset needs to be constructed for model training. This dataset should contain diverse images taken under different lighting conditions, angle changes, and weather conditions, and special attention should be paid to collecting various obstacles that may appear in the fire passage, such as debris, people, vehicles, etc. Each picture needs to be manually or semi-automatically annotated to clearly mark the position of the fire passage, whether there are objects occupying the passage, the category, size, and precise position of the occupying objects.
[0030] Next, a suitable large vision model architecture is selected for training. In one embodiment, the trained large vision model used is the RT-DETR model based on the Transformer architecture. This architecture has been proven to have excellent object detection capabilities, especially not being easily overfitted when dealing with large-scale data. In addition, an important feature of the RT-DETR model is that it can achieve open-vocabulary object detection, that is, for new categories that do not appear in the training set, the model can also effectively identify them. This is achieved by modifying the weights of the object classification layer at the end of the network: using CLIP (Contrastive Language–Image Pre-training) to encode the categories and replacing the original weights of the object classification layer with the embedding results. In this way, when encountering new categories during the application stage, only the corresponding text description needs to be provided without retraining the entire model.
[0031] Once the model is trained, it can be applied to actual monitoring tasks. The system receives a real-time video stream from a camera and extracts a series of key frames as input to the visual large model. The selection of these key frames is based on specific criteria, such as when there are significant changes in the scene (e.g., a new object enters the field of view or the position of an existing object changes), or sampling is performed at fixed time intervals. For each key frame, the visual large model outputs detailed detection results, including whether the fire escape is unobstructed and whether there are any obstacles; marks the specific coordinates of all objects appearing in the fire escape, which helps to understand the relative position of the objects with respect to the channel boundary; analyzes the movement trajectories of the objects in the time series to determine whether they are temporarily parked or occupy the channel for a long time. For example, if an object appears in one key frame but disappears in the following few frames, it may be a temporary stop; conversely, if the object persists in multiple consecutive key frames, there may be a risk of long-term occupation.
[0032] In particular, considering the relatively low reliability of single-frame results, this application proposes a method based on sequence analysis and result recombination for more in-depth analysis. Specifically, the number of key frames within a certain time period is designed, the time-series data is evenly sampled, each frame of video image is detected, and matrix analysis is performed on the detection results. First, calculate whether each target is within the area, and for the targets within the area, calculate the proportion of the fire area they occupy and summarize it. Then, perform time-series position analysis on individual targets to determine whether there is movement and assign a label. This method not only improves the accuracy of detection but also enhances the understanding of dynamic changes.
[0033] By continuously monitoring and recording this information, this application can construct a time series regarding the occupancy of fire lanes. This time series is not just a series of static results, but rather a record of a dynamic process, reflecting the behavioral patterns of objects in the fire lane and their development trends over time. For example, if an object is found to appear in the fire lane starting from a key frame and remains unchanged in the following multiple frames, this may indicate that the object has occupied the lane for some time and its potential safety risks need further attention. To make the detection results more intuitive and understandable, this information can also be converted into concise text descriptions. For example, "At the 5th second, the entrance of the fire lane is occupied by a pile of sundries, the relative position of the object is (x1, y1, w, h), and the occupancy ratio is s". Here, the occupancy ratio s is obtained by calculating the ratio of the target occupancy area to the entire fire lane area, which is different from the intersection over union (IoU) concept in traditional object detection. In this application, for example, if the other category area is s1 and the target area is s2, then first calculate the intersection area s', and then calculate the ratio of s' to s2. After the entire sequence runs, a set T = {text1, text2, text3,...} composed of multiple such texts will be generated and used as input together with a specially designed prompt for the large language model to perform the final warning level analysis.
[0034] Exemplarily, in step S3, after adding a prompt at the tail of the time series of the detection results, it is input into the fire lane occupancy warning level analysis engine based on the large language model to obtain the fire lane occupancy warning level analysis result. It should be understood that although the vision large model can provide accurate object detection information, it may be insufficient for decision-making support in complex scenarios (such as judging whether there is occupancy, the duration, and a reasonable warning level). The large language model has powerful natural language understanding and generation capabilities. Through appropriate prompt guidance, it can better understand the meaning represented by the detection results and then perform higher-level analysis. Specifically, the prompt provides clear task guidance for the large language model, that is, to judge whether the fire lane is occupied based on the provided description and output detailed warning information. This method not only improves the accuracy of the final analysis result but also makes the whole process more automated and intelligent.
[0035] That is, in the process of adding a prompt word to the tail of the time series of the detection results and then inputting it into the fire lane occupancy warning level analysis engine based on a large language model to obtain the fire lane occupancy warning level analysis result, by combining the time series of the detection results with the prompt word, it can help the large language model better understand the dynamic changes and their meanings of the fire lane status in the time series data of the detection results. For example, information such as the position movement, appearance, or disappearance of an object in consecutive frames can be conveyed to the model in this way, enabling it to more accurately analyze the meanings behind these changes, thereby obtaining a more accurate fire lane occupancy warning level analysis result, which helps to provide a more reasonable decision-making.
[0036] In one embodiment, as Figure 3 shown, after adding a prompt word to the tail of the time series of the detection results and inputting it into the fire lane occupancy warning level analysis engine based on a large language model to obtain the fire lane occupancy warning level analysis result, it includes: S31, performing semantic embedding encoding on each detection result in the time series of the detection results to obtain a time series of detection result semantic embedding encoding vectors; S32, performing context semantic association encoding on the time series of the detection result semantic embedding encoding vectors to obtain a fire lane detection result time series context semantic association encoding vector; S33, adding the prompt word to the tail of the fire lane detection result time series context semantic association encoding vector.
[0037] In the step S31, each detection result in the time series of the detection results is subjected to semantic embedding encoding to obtain a time series of detection result semantic embedding encoding vectors. Considering that the original detection results (such as object position, category, and occupancy ratio, etc.) exist in the form of structured data or text descriptions. In order to enable these information to be effectively processed by the large language model, they need to be transformed into a unified mathematical expression form. Therefore, in the technical solution of this application, each detection result in the time series of the detection results is subjected to semantic embedding encoding to obtain a time series of detection result semantic embedding encoding vectors. Through the method of semantic embedding encoding, different types of detection result data types can be transformed into the same data expression form, and at the same time, the implicit association relationships between different data types in each detection result can be captured, such as the relationship between the object position, category, and occupancy ratio of an obstacle, which enhances the understanding ability of the model, thereby helping to perform level analysis and warning on the occupancy situation of the fire lane.
[0038] In one embodiment, a pre-trained language model, such as a BERT or GPT series model, can be used to encode each piece of descriptive text. This process involves splitting the text into words or sub-word units and then mapping each unit to a point in a multi-dimensional space according to the weight parameters learned within the model. The advantage of this is that it can capture the semantic relationships between words and also have a certain understanding ability even for newly emerging concepts that are not covered in the training set.
[0039] In step S32, a context semantic association encoding is performed on the time series of the detection result semantic embedding encoding vectors to obtain a fire lane detection result time series context semantic association encoding vector. It should be understood that since each detection result semantic embedding encoding vector in the time series of the detection result semantic embedding encoding vectors respectively contains the performance feature information regarding the fire lane status in each key frame of the fire lane status monitoring, including embedded semantic feature information such as object position, category, and occupancy ratio, these detection result semantic embedding encoding vectors can be regarded as a compact and meaningful representation of the fire lane status at that moment. When detecting the occupancy of the fire lane based on the fire lane status shown in a single-frame image, it is necessary to artificially set multiple safety levels in advance, such as which type of target occupancy is considered a high risk, etc., based on human experience. This easily leads to misjudgment problems due to the unreliability of the single-frame analysis result. And there is no need to set the warning level scenario according to human experience, such as vehicle occupancy being regarded as a high-risk situation. Also, considering that there are temporal edge change features and correlation relationships in the time dimension between the fire lane status features in each key frame of the fire lane status monitoring, this correlation information reflects the temporal change situation of information such as the position of obstacles in the fire lane, which helps to more accurately capture how the objects in the fire lane move or whether the occupancy situation is deteriorating, providing a basis for the subsequent analysis of the fire lane occupancy warning level. Based on this, in order to better understand and predict the occupancy status and change trend of the fire lane, in the technical solution of this application, a context semantic association encoding is further performed on the time series of the detection result semantic embedding encoding vectors to obtain a fire lane detection result time series context semantic association encoding vector. In particular, the process of context semantic association encoding is executed through a feature message passing network, which explicitly measures the propagation efficiency of the fire lane status features under each key frame by introducing the feature jump degree and the message passing space span, so as to highlight important time node information and dynamically adjust the information flow during the propagation process of the fire lane status information to obtain a fire lane detection result time series context semantic association feature message passing expression with good generalization ability and expressiveness.
[0040] In one embodiment, as Figure 4As shown, in step S32, context semantic association encoding is performed on the time series of the semantic embedding encoding vectors of the detection results to obtain a time series context semantic association encoding vector for fire lane detection results, including: S321, calculating the feature jump degree and message passing space span of each semantic embedding encoding vector in the time series of the semantic embedding encoding vectors of the detection results to obtain a time series of detection result semantic feature jump degrees and a time series of detection result semantic feature message passing space spans; S322, based on the time series of the detection result semantic feature jump degrees and the time series of the detection result semantic feature message passing space spans, performing significant weighted aggregation on the time series of the semantic embedding encoding vectors of the detection results to obtain the time series context semantic association encoding vector for fire lane detection results.
[0041] In one embodiment, as Figure 5 shown, in step S321, calculating the feature jump degree and message passing space span of each semantic embedding encoding vector in the time series of the semantic embedding encoding vectors of the detection results to obtain a time series of detection result semantic feature jump degrees and a time series of detection result semantic feature message passing space spans, including: S3211, calculating the feature jump degree of each semantic embedding encoding vector of the detection results to obtain a time series of the detection result semantic feature jump degrees. Specifically, this process can be represented by the formula:
[0042]
[0043]
[0044]
[0045] where is the time series of the semantic embedding encoding vectors of the detection results, , , and are respectively the 1st, 2nd, th, and th semantic embedding encoding vectors in the time series of the semantic embedding encoding vectors of the detection results, is the eigenvalue at the th position in , is the number of eigenvalues in and are respectively the th and th detection result semantic weighted average eigenvalues in the time series of the detection result semantic weighted average eigenvalues, is the jump degree of the semantic features of the detection result in the time series at the th detection result semantic feature jump degree.
[0046] S3212. Calculate the number of message transmissions between each detection result semantic embedding coding vector and the last detection result semantic embedding coding vector in the time series of the detection result semantic embedding coding vectors to obtain a sequence of detection result semantic transmission times; S3213. Based on the sequence of the detection result semantic transmission times, calculate the message transmission space span of each detection result semantic embedding coding vector to obtain the time series of the detection result semantic feature message transmission space span. Specifically, this process can be expressed by the formula:
[0047]
[0048] where represents to the number of message transmissions, is the jump degree of the semantic features of the detection result in the time series at the th detection result semantic feature message transmission space span.
[0049] Specifically, by calculating the feature jump degree of each detection result semantic embedding coding vector, the degree of scene change between different time points can be quantified. The so-called feature jump degree refers to the difference size between the semantic embedding coding vectors at two or more adjacent time points. In a specific example, if an object suddenly appears in the fire passage and occupies a part of the space, then the corresponding semantic embedding coding vector will change significantly, which will be reflected in a higher feature jump degree value. On the contrary, if the scene is relatively stable and there is no obvious object movement or addition, the feature jump degree will be lower. Therefore, the generated time series of the detection result semantic feature jump degree can help us identify which time periods important events have occurred, such as the appearance of new obstacles or significant changes in the positions of existing objects.
[0050] Next, calculate the number of message transmissions between each encoding vector in the time series of the detected result semantic embedding encoding vectors and the last detected result semantic embedding encoding vector. In fact, it is measuring the information update frequency at each time point. The number of message transmissions here refers to how many times new information is introduced from the current time point to the end of the sequence. In a practical application scenario, assume that a camera in an underground garage takes a picture every few seconds and sends it to a visual large model for processing. When a vehicle enters the fire lane and stays there, the first few key frames may record this new situation. However, as the vehicle continues to park, subsequent key frames may no longer contain additional new information. At this time, the number of message transmissions can help distinguish those time points that truly bring new changes (such as the vehicle first entering) from other time points that just repeatedly confirm known situations.
[0051] Finally, based on the sequence of the number of message transmissions obtained above, further calculate the message transmission spatial span of each detected result semantic embedding encoding vector, aiming to evaluate the influence range of the information carried at each time point in the entire time series. The message transmission spatial span reflects how the information in a certain time period spreads and affects the state descriptions in subsequent time periods. In a typical scenario, if a private car parks in the fire lane for a long time, the first few detection results will record this fact. As time goes by, even if the vehicle's position does not change significantly, it still has a continuous impact on the entire time series. This impact can be reflected by the message transmission spatial span: the early key frames have a larger spatial span because they define the basic conditions for all subsequent time points; while the later key frames gradually reduce their influence, unless there are new changes (such as the vehicle leaving). Therefore, the constructed time series of the message transmission spatial span of the detected result semantic features provides important clues about the information propagation path and intensity, enabling us to more accurately judge which factors play a key role in the overall trend.
[0052] In one embodiment, as Figure 6 shown, in the step S322, based on the time series of the detected result semantic feature jump degrees and the time series of the message transmission spatial spans of the detected result semantic features, perform significant weighted aggregation on the time series of the detected result semantic embedding encoding vectors to obtain the fire lane detection result time series context semantic association encoding vectors, including: S3221, based on the time series of the detected result semantic feature jump degrees and the time series of the message transmission spatial spans of the detected result semantic features, calculate the message transmission significant weights of each detected result semantic embedding encoding vector in the time series of the detected result semantic embedding encoding vectors to obtain the time series of the detected result semantic feature message transmission significant weights. Specifically, this process can be represented by the formula as:
[0053]
[0054]
[0055] Among them, and are modulation parameters, is the th detection result semantic feature message passing significant factor in the time series of detection result semantic feature message passing significant factors, is a normalization function, is a masking operation, is a predetermined threshold, is the th detection result semantic feature message passing significant weight in the time series of detection result semantic feature message passing significant weights.
[0056] S3222, using the time series of the detection result semantic feature message passing significant weights, calculate the weighted sum of the time series of the detection result semantic embedding coding vectors to obtain the fire lane detection result time series context semantic association coding vector. Specifically, this process can be expressed by the formula as:
[0057]
[0058] Among them, is the number of vectors in the time series of the detection result semantic embedding coding vectors, is the fire lane detection result time series context semantic association coding vector.
[0059] Specifically, by combining the time series of the detection result semantic feature jump degree (i.e., the degree of scene change between adjacent time points) and the time series of the detection result semantic feature message passing spatial span (i.e., the description of how the information in a certain time period spreads and affects the subsequent time period state), the importance of each time point can be evaluated more comprehensively. A high jump degree means that the scene has changed greatly, which may indicate the occurrence of important events; a larger message passing spatial span means that the information at this time point has a wide impact on the subsequent state. Considering these two factors comprehensively, a more reasonable message passing significant weight can be assigned to each time point, so as to ensure that those time points that truly bring new changes or have long-term impacts are given sufficient attention.
[0060] Finally, using the calculated significant message passing weights as coefficients, perform a weighted sum on the time series of the semantic embedding encoding vectors of the entire detection result. In this way, a time-series context semantic association encoding vector for the fire lane detection result can be obtained. This encoding vector is not just a simple numerical summary but rather the result of integrating information from all time points and their interrelationships. It retains the key features regarding the changes in the fire lane status in the original data and highlights the factors that play important roles in the overall trend. For example, in the above example, if the system detects that a private car has occupied the fire lane for a long time, it can not only accurately report the specific time and location of the occupation but also predict the impact of this situation on future traffic capacity based on the time-series context semantic association encoding vector. This prediction function enables the system to identify potential risks in advance and promptly trigger high-priority alerts to notify relevant personnel to take actions to move the vehicle, thus avoiding the occurrence of dangerous situations.
[0061] In one embodiment, based on the time series of the jump degrees of the semantic features of the detection result and the time series of the message passing spatial spans of the semantic features of the detection result, calculating the message passing significant weights of each detection result semantic embedding encoding vector in the time series of the detection result semantic embedding encoding vectors to obtain a time series of the message passing significant weights of the semantic features of the detection result includes: calculating the message passing significant factors of each detection result semantic embedding encoding vector based on the feature jump degrees and message passing spatial spans of each detection result semantic embedding encoding vector in the time series of the detection result semantic embedding encoding vectors to obtain a time series of the message passing significant factors of the semantic features of the detection result; performing gated screening on the time series of the message passing significant factors of the semantic features of the detection result to obtain the time series of the message passing significant weights of the semantic features of the detection result.
[0062] In one embodiment, calculating a message passing significance factor for each of the detection result semantic embedding encoding vectors in the time series of the detection result semantic embedding encoding vectors based on the feature jump degree and the message passing space span of each of the detection result semantic embedding encoding vectors in the time series to obtain a time series of the detection result semantic feature message passing significance factor includes: extracting the detection result semantic feature jump degree and the detection result semantic feature message passing space span corresponding to a first detection result semantic embedding encoding vector from the time series of the detection result semantic feature jump degree and the time series of the detection result semantic feature message passing space span; calculating the square of the detection result semantic feature jump degree corresponding to the first detection result semantic embedding encoding vector to obtain a detection result semantic feature jump degree square modulation representation value; and performing weighted modulation optimization representation on the detection result semantic feature jump degree square modulation representation value by using the detection result semantic feature message passing space span corresponding to the first detection result semantic embedding encoding vector as a modulation coefficient to obtain the detection result semantic feature message passing significance factor.
[0063] In summary, the process of performing context semantic association encoding on the time series of the detection result semantic embedding encoding vectors can, by combining the feature jump degree and the space span, understand the broader temporal connectivity and long-distance temporal dependence relationships of the fire lane state while maintaining the static fire lane state feature representation at each moment. This approach integrates the dynamic change pattern of the fire lane state in the time dimension, enabling subsequent large language models or other decision-making modules to make accurate early warning level judgments based on richer and more comprehensive data. In particular, the process of context semantic association encoding automatically identifies the fire lane states at these key time nodes through a feature message passing network by calculating the message passing significance factor and assigns higher weights to ensure that the model can focus on the most relevant fire lane occupancy information for early warning level analysis. This method is particularly suitable for application scenarios that require real-time response and accurate judgment, such as fire lane occupancy early warning. It can effectively capture the development trend of occupancy behavior, timely detect potential risks, and provide a scientific basis for management personnel so that they can take appropriate actions to ensure public safety.
[0064] In the step S33, the prompt word is added to the tail of the time-series context semantic association encoding vector of the fire lane detection result. It should be understood that through the above processing, the time-series context semantic association encoding vector of the fire lane detection result has been obtained. Even with the time-series context semantic association encoding vector of the fire lane detection result, the large language model still needs clear task guidance to effectively perform reasoning and decision-making. The role of the prompt word is to provide specific instructions to the large language model, telling it what to do next and how to interpret the input data. Specifically, in the prompt word, the task objective is first briefly described, that is, to analyze whether the fire lane is occupied. At the same time, the model is required to judge the warning level according to the occupancy situation (such as occupancy time, object type). For example, if the occupancy lasts for a long time and it is an object that affects passage (such as a person), it can be judged as a high warning. Further, further restrictions can be added to the prompt word to generate a more standardized analysis result. For example, the prompt word is "Please judge whether there is occupancy in the fire lane according to the above description and give the following information: a. Whether there is occupancy (yes / no). b. If there is occupancy, what is the duration (in seconds). c. Give a reasonable warning level (low / medium / high) according to the occupancy duration and the type of the occupied object (for example, sundries or people)".
[0065] Here, in the case where each detection result semantic embedding encoding vector in the time series of the detection result semantic embedding encoding vector respectively represents the low-dimensional encoded semantic features of the corresponding detection result, when performing feature sequence message passing based on the feature jump degree, the generation of insufficient global semantic space feature jump degree correlation correspondence caused by the local semantic space difference of the encoded semantics will cause the sparse sequence message association transmission of the time-series context semantic association encoding vector of the fire lane detection result, thereby reducing the accuracy of the fire lane occupancy warning level analysis result obtained by the fire lane occupancy warning level analysis engine based on the large language model due to the lack of generated inference degree.
[0066] In a preferred example, after adding the prompt word to the tail of the time-series context semantic association encoding vector of the fire lane detection result, inputting it into the fire lane occupancy warning level analysis engine based on the large language model to obtain the fire lane occupancy warning level analysis result includes:
[0067] Determine the number of hyper-distribution eigenvalue numbers in the time-series context semantic association encoding vector of the fire lane detection result that are greater than the threshold difference from the feature mean, and calculate the reciprocal of the logarithm to the base 2 of the number of hyper-distribution eigenvalue numbers and the exponential value to the base of the natural constant of the reciprocal of the number of hyper-distribution eigenvalue numbers respectively to obtain the first fire lane detection result time-series context semantic association response field cell representation value The value of the cell representation of the temporal context semantic association response field for the second fire escape detection result :
[0068]
[0069]
[0070]
[0071] Wherein is the number of hyper-distribution eigenvalue in the temporal context semantic association coding vector of the fire escape detection result, and is the threshold hyperparameter,[[]] represents the temporal context semantic association coding vector of the fire escape detection result,[[]] represents the th eigenvalue of the temporal context semantic association coding vector of the fire escape detection result,[[]] represents the calculation of the number of,[[]] represents the mean value of all eigenvalues of the temporal context semantic association coding vector of the fire escape detection result;[[]]
[0072] Calculate the hyperbolic sine function value of the sum of squares of all eigenvalues of the temporal context semantic association coding vector of the fire escape detection result:[[]]
[0073]
[0074] Wherein,[[]] represents the length of the temporal context semantic association coding vector of the fire escape detection result,[[]] represents the hyperbolic sine function,[[]] represents the hyperbolic sine function value;[[]]
[0075] And calculate the exponential value of the hyperbolic sine function value with the natural constant as the base and then divide it by the square of the length of the temporal context semantic association coding vector of the fire escape detection result to obtain the multi-dimensional field manifold value of the temporal context semantic association response of the fire escape detection result[[]] , wherein,[[]] represents the natural constant;[[]]
[0076] Calculate the multi-dimensional field manifold value of the temporal context semantic association response of the fire escape detection result[[]] relative to the value of the cell representation of the temporal context semantic association response field of the first fire escape detection result[[]] and the value of the cell representation of the temporal context semantic association response field of the second fire escape detection result[[]] The covariance integration value of the temporal context semantic association response of the fire escape detection result:[[]]
[0077]
[0078] Among them, represents the covariance integration value of the temporal context semantic association response of the fire lane detection result;
[0079] Calculate the power function eigenvector of the temporal context semantic association coding vector of the fire lane detection result with the reciprocal of the covariance integration value of the temporal context semantic association response of the fire lane detection result, and calculate the autocorrelation matrix of the power function eigenvector to obtain the adaptive representation matrix of the temporal context semantic association response of the fire lane detection result, that is:
[0080]
[0081] Among them, represents calculating the power function eigenvector of the temporal context semantic association coding vector of the fire lane detection result with the reciprocal of the covariance integration value of the temporal context semantic association response of the fire lane detection result, represents the transpose of the vector, represents matrix multiplication, represents the adaptive representation matrix of the temporal context semantic association response of the fire lane detection result;
[0082] Multiply the temporal context semantic association coding vector of the fire lane detection result and the adaptive representation matrix of the temporal context semantic association response of the fire lane detection result to obtain an optimized temporal context semantic association coding vector of the fire lane detection result, that is:
[0083]
[0084] Here, the temporal context semantic association coding vector of the fire lane detection result is a row vector, represents the optimized temporal context semantic association coding vector of the fire lane detection result;
[0085] After adding a prompt word to the tail of the optimized temporal context semantic association coding vector of the fire lane detection result, input it into the fire lane occupancy warning level analysis engine based on the large language model to obtain the fire lane occupancy warning level analysis result.
[0086] That is, for the attribute cluster of the temporal context semantic association coding vector of the fire passage detection result, due to the lack of the global adaptive mapping mechanism caused by the cross-distribution field association under the discrete arrangement of preset parameters, a multi-dimensional tensor field cell representation framework is constructed with the cross-covariance coupling driving scheme of the temporal context semantic association coding vector of the fire passage detection result, and the deep interaction law of its parameter space is revealed by analyzing the fully connected topological association composite mode. Thus, by adopting the gradient-based multi-dimensional field implicit factor simulation method, a non-linear association modeling system of the parameter space is established, so as to realize the topological reconstruction of the probabilistic trajectory evolution mode of the temporal context semantic association coding vector of the fire passage detection result. In this way, the accuracy of the fire passage occupancy warning level analysis result obtained by improving the input of the temporal context semantic association coding vector of the fire passage detection result into the fire passage occupancy warning level analysis engine based on the large language model is improved.
[0087] In summary, the method for analyzing the fire passage occupancy warning level based on the large model according to the embodiment of the present application is elucidated. It obtains the time series of the key frames of the fire passage status monitoring by real-time monitoring and collecting the status of the fire passage through a camera, and then uses the vision large model to detect these key frames of the fire passage status monitoring to identify the fire passage status, object position and change situation in each key frame, so as to generate the corresponding detection results. Finally, the large language model and the prompt information are used to analyze the fire passage occupancy warning level of these analyzed detection results to obtain the analysis result. Compared with the traditional single-frame object detection analysis result, this helps to improve the speed and reliability of the fire emergency response, thus reducing the safety hazards caused by the blockage of the fire passage.
[0088] Figure 7 It is a schematic block diagram of the fire passage occupancy warning level analysis system based on the large model according to the embodiment of the present application. As Figure 7As shown in the figure, the large model-based fire lane occupancy warning level analysis system 100 includes: a fire lane status monitoring and acquisition module 110, which is used to obtain the time series of key frames of fire lane status monitoring collected by a camera; a fire lane status detection module 120, which is used to use the trained vision large model to detect each key frame of the fire lane status monitoring in the time series of key frames of fire lane status monitoring to obtain a time series of detection results, and the detection results include fire lane status, object position, and change situation; a fire lane occupancy warning level analysis module 130, which is used to add a prompt word to the tail of the time series of detection results and then input it into the fire lane occupancy warning level analysis engine based on the large language model to obtain the fire lane occupancy warning level analysis result, and the prompt word is "Please judge whether there is occupancy in the fire lane according to the above description and give the following information: a. Whether there is occupancy; b. If there is occupancy, what is the duration; c. Give a reasonable warning level according to the occupancy duration and the type of occupied object".
[0089] In one embodiment, the trained vision large model is an RT-DETR model based on the transformer architecture.
[0090] Here, those skilled in the art can understand that the specific operations of each module and unit in the above large model-based fire lane occupancy warning level analysis system have been introduced in detail in the description of the large model-based fire lane occupancy warning level analysis method above, and therefore, the repeated description thereof will be omitted. Figures 1 to 6 The description of the method based on the large model of the fire lane occupancy warning level analysis method has been introduced in detail, and therefore, the repeated description thereof will be omitted.
[0091] The embodiment of the present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer is enabled to implement the methods in the above embodiments of the present application.
[0092] The embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions run on a computer, the computer is enabled to implement the methods in the above embodiments of the present application.
[0093] The embodiment of the present application also provides a chip, which includes a circuit for executing the methods in the above embodiments of the present application.
[0094] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0095] In the embodiments of the present application, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. The use of prefix words such as ordinal numbers for distinguishing described objects in the embodiments of the present application does not constitute a restriction on the described objects. The statements of the described objects refer to the descriptions in the context of the claims or embodiments, and should not constitute unnecessary restrictions due to the use of such prefix words.
[0096] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0097] In each embodiment of the present application, if there is no special description and logical conflict, the terms and / or descriptions between the embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0098] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0099] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0100] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A fire passage occupancy warning level analysis method based on a large model, characterized in that: include: Obtain the time series of key frames of fire passage status monitoring collected by the camera; Using the trained visual big model, each fire passage state monitoring key frame in the time series of the fire passage state monitoring key frames is detected to obtain a time series of detection results, wherein the detection results include the fire passage state, object position and change; After adding the prompt word at the end of the time series of the detection result, it is input into the fire passage occupancy warning level analysis engine based on the large language model to obtain the fire passage occupancy warning level analysis result, and the prompt word is "Please judge whether the fire passage is occupied according to the above description, and give the following information:
1. Whether it is occupied; 2. If there was occupation, what was the duration? 3. Give a reasonable level of warning, depending on the duration of occupation and the type of occupied object”; The process of performing context semantic association coding on the time series of the detection result semantic embedding coding vector to obtain the fire channel detection result time series context semantic association coding vector comprises: The feature jump degree of the semantic embedding coding vector of each detection result is calculated to obtain the time series of the semantic feature jump degree of the detection result. The calculation process can be expressed as follows: ; in, is the time series of the semantic embedding encoding vector of the detection result, and are the first, second, and third time series of the semantic embedding encoding vector of the detection result. and The detection result semantic embedding encoding vector, yes Middle The eigenvalues at the positions, yes The number of eigenvalues in , and They are the time series of the semantic weighted average feature values of the detection results. and The semantically weighted average feature value of the detection results, is the time series of the semantic feature jump degree of the detection result. The semantic feature jump degree of each detection result; Calculating the number of message transmissions between each detection result semantic embedding coding vector and the last detection result semantic embedding coding vector in the time series of the detection result semantic embedding coding vector to obtain a sequence of detection result semantic transmission times; Based on the sequence of the number of times the semantics of the detection results are transmitted, the message transmission space span of the semantic embedding coding vectors of the respective detection results is calculated to obtain a time series of the message transmission space span of the semantic features of the detection results; Based on the time series of the jump degree of the semantic features of the detection results and the time series of the spatial span of the semantic feature message transmission of the detection results, the time series of the semantic embedding coding vector of the detection results is significantly weighted and aggregated to obtain the temporal context semantic association coding vector of the fire escape detection result.
2. The fire passage occupancy warning level analysis method based on a large model according to claim 1 is characterized in that: The trained visual model is a RT-DETR model based on the transformer architecture.
3. The fire passage occupancy warning level analysis method based on a large model according to claim 2 is characterized in that: After adding the prompt word at the end of the time series of the detection result, it is input into the fire passage occupancy warning level analysis engine based on the large language model to obtain the fire passage occupancy warning level analysis result, including: Performing semantic embedding coding on each detection result in the time series of the detection results to obtain a time series of semantic embedding coding vectors of the detection results; Performing context semantic association coding on the time series of the detection result semantic embedding coding vector to obtain a fire channel detection result time series context semantic association coding vector; The prompt word is added to the end of the temporal context semantic association encoding vector of the fire passage detection result.
4. The fire passage occupancy warning level analysis method based on a large model according to claim 3 is characterized in that: Based on the time series of the jump degree of the semantic feature of the detection result and the time series of the spatial span of the semantic feature message transmission of the detection result, the time series of the semantic embedding coding vector of the detection result is significantly weighted and aggregated to obtain the temporal context semantic association coding vector of the fire escape detection result, including: Based on the time series of the jump degree of the semantic feature of the detection result and the time series of the message transmission space span of the semantic feature of the detection result, calculating the message transmission significance weight of each detection result semantic embedding coding vector in the time series of the semantic embedding coding vector of the detection result to obtain the time series of the message transmission significance weight of the semantic feature of the detection result; The weighted sum of the time series of the semantic embedding coding vectors of the detection results is calculated based on the time series of the significant weights of the semantic feature messages of the detection results to obtain the temporal context semantic association coding vector of the fire escape detection result.
5. The fire passage occupancy warning level analysis method based on a large model according to claim 4 is characterized in that: Based on the time series of the jump degree of the semantic feature of the detection result and the time series of the message transmission space span of the semantic feature of the detection result, calculating the message transmission significance weight of each detection result semantic embedding coding vector in the time series of the semantic embedding coding vector of the detection result to obtain the time series of the message transmission significance weight of the semantic feature of the detection result, including: Based on the feature jump degree and message transfer space span of each detection result semantic embedding coding vector in the time series of the detection result semantic embedding coding vector, calculating the message transfer significance factor of each detection result semantic embedding coding vector to obtain the time series of the detection result semantic feature message transfer significance factor; The time series of the semantic feature message transmission significance factor of the detection result is gated and screened to obtain the time series of the semantic feature message transmission significance weight of the detection result.
6. The fire passage occupancy warning level analysis method based on a large model according to claim 5 is characterized in that: Based on the feature jump degree and message transfer space span of each detection result semantic embedding coding vector in the time series of the detection result semantic embedding coding vector, calculating the message transfer significance factor of each detection result semantic embedding coding vector to obtain the time series of the detection result semantic feature message transfer significance factor, including: Extracting the detection result semantic feature jump degree and the detection result semantic feature message transfer space span corresponding to the first detection result semantic embedding coding vector from the time series of the detection result semantic feature jump degree and the detection result semantic feature message transfer space span; Calculating the square of the jump degree of the semantic feature of the detection result corresponding to the first detection result semantic embedding coding vector to obtain a square modulation representation value of the jump degree of the semantic feature of the detection result; The detection result semantic feature message transmission space span corresponding to the first detection result semantic embedding coding vector is used as the modulation coefficient to perform weighted modulation optimization representation on the detection result semantic feature jump degree square modulation representation value to obtain the detection result semantic feature message transmission significance factor.
7. A large model-based fire passage occupancy warning level analysis system, used to execute the large model-based fire passage occupancy warning level analysis method according to claim 1, characterized in that: include: The fire passage status monitoring and acquisition module is used to obtain the time series of the fire passage status monitoring key frames collected by the camera; A fire passage state detection module, used to detect each fire passage state monitoring key frame in the time series of the fire passage state monitoring key frames using the trained visual large model to obtain a time series of detection results, wherein the detection results include the fire passage state, object position and change; The fire passage occupation warning level analysis module is used to add a prompt word at the end of the time series of the detection result, and then input it into the fire passage occupation warning level analysis engine based on the large language model to obtain the fire passage occupation warning level analysis result, wherein the prompt word is "Please judge whether the fire passage is occupied according to the above description, and give the following information:
1. Whether it is occupied; 2. If there was occupation, what was the duration? 3. Give reasonable warning levels, depending on the duration of occupancy and the type of occupied object”.
8. The fire passage occupancy warning level analysis system based on a large model according to claim 7 is characterized in that: The trained visual model is a RT-DETR model based on the transformer architecture.
Citation Information
Patent Citations
Early warning method and device based on multi-modal large model, equipment and medium
CN119206578A
Illegal occupation system for fire fighting access
CN119296035A