In-station escalator safety monitoring method and system, processing equipment and storage medium
By combining the SAM image segmentation model and the Mind Chain multimodal large model, the problem of insufficient spatial position perception of the multimodal large model in escalator safety monitoring is solved, realizing high-precision abnormal behavior detection in the escalator area and improving the accuracy and robustness of monitoring.
Patent Information
- Application Number
- CN202511621543.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-07
AI Technical Summary
In existing technologies, multimodal large models have a weak ability to perceive spatial location information in escalator safety monitoring, leading to frequent misjudgments. Furthermore, traditional single Prompt is difficult to adapt to the changing escalator scenarios, affecting the accuracy and reliability of detection.
The SAM image segmentation model is used for preprocessing to identify the spatial boundaries between escalators and stairs. Combined with the multimodal large model of the thought chain, prompt words are adaptively matched. By segmenting mask images and analyzing scene features, abnormal behavior in the escalator area is accurately identified.
It significantly improves the regional positioning accuracy and overall reliability of abnormal event detection, realizes efficient and accurate monitoring in different scenarios, reduces false alarms, and improves the robustness and adaptability of the model.
Smart Images

Figure CN121074809A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of safety monitoring, in particular to an escalator safety monitoring method and system in a station, a processing device and a storage medium. BACKGROUND
[0002] As an important part of modern urban public transportation infrastructure, escalators have been widely used in railway stations, shopping malls, subway stations, airports and other public places, and bear the key function of large passenger flow transportation. However, due to the complex situations such as dense passenger flow in the escalator area, various types of luggage carried by passengers, and unpredictable passenger behavior, the frequency of escalator falling incidents is high, which brings serious safety hazards to passengers. Therefore, how to efficiently and accurately monitor the safety of the escalator is imminent.
[0003] With the development of artificial intelligence, some enterprises use video monitoring and large model deep semantic understanding analysis to monitor the behavior of personnel on the escalator in real time and issue abnormal alarms. However, due to the different positions and angles of the escalator monitoring cameras, the different models of the cameras, and the great differences in the escalator scenes displayed by different cameras, the same monitoring method cannot be effectively generalized in different scenes. Currently, multi-modal large models have significant limitations in abnormal event detection tasks, and their spatial position information perception ability is weak, which cannot effectively distinguish events occurring in different areas. In the escalator safety monitoring scene, the monitoring screen usually contains both the escalator and the staircase area. When abnormal situations such as falling occur in the staircase area, the multi-modal large model will misjudge it as an escalator area event, causing the alarm system to be frequently triggered, which seriously affects the accuracy and reliability of the detection.
[0004] In addition, in the escalator monitoring scene, due to the differences in camera models, the resolution of the collected images presents diversity. High-resolution images carry rich detailed information, while low-resolution images have a feature information missing problem. At the same time, due to different scene configurations such as single escalator operation and multi-escalator operation in the escalator area, it is difficult for traditional single Prompt (prompt word, instruction of large model) to adapt to complex and variable detection requirements, which seriously restricts the generalization ability and accuracy of abnormal event detection. Therefore, how to design and build a Prompt system with scene universality has also become a problem that needs to be solved in the industry. SUMMARY
[0005] To solve the above problems, the purpose of the present application is to provide an escalator safety monitoring method, system, processing device and storage medium with high detection accuracy, high reliability and scene universality.
[0006] To achieve the above purpose, the present application adopts the following technical solutions: in the first aspect, an escalator safety monitoring method in a station is provided, comprising: The monitoring video of the escalator in the station is acquired, and a SAM image segmentation model is adopted to generate a segmentation mask picture based on an input prompt word set and a mask feature map, so as to determine the escalator related area corresponding to each monitoring scene picture in the monitoring video and obtain the position information of each escalator related area relative to the corresponding monitoring scene picture. A multi-modal large model based on a thinking chain is adopted to perform prompt word adaptive matching on the monitoring scene picture, and core detection is performed according to the generated segmentation mask picture and the matched prompt word, so as to determine whether a high-risk abnormal behavior event such as falling of a person occurs in the escalator area in the station. When a high-risk abnormal behavior event occurs in the escalator area in the station, the alarm position and alarm level of the escalator in the station are determined according to the alarm result of the multi-modal large model based on the thinking chain and the position information of each escalator related area relative to the corresponding monitoring scene picture.
[0007] Further, the acquisition of the monitoring video of the escalator in the station and the adoption of the SAM image segmentation model to generate the segmentation mask picture based on the input prompt word set and the mask feature map to determine the escalator related area corresponding to each monitoring scene picture in the monitoring video and obtain the position information of each escalator related area relative to the corresponding monitoring scene picture comprises: Acquiring the monitoring video of the escalator in the station; Frame extraction is performed on the acquired monitoring video to obtain corresponding monitoring scene pictures; Each monitoring scene picture is input into an image encoder of the SAM image segmentation model for processing, and each monitoring scene picture is converted into an image embedding; The text prompt word is input into a Prompt encoder of the SAM image segmentation model for processing, and the text prompt word is converted into a prompt embedding; The mask feature map is input into a down-sampling module of the SAM image segmentation model for processing to obtain a down-sampled mask feature map; The converted image embedding and prompt embedding and the down-sampled mask feature map are input into a lightweight mask decoder of the SAM image segmentation model for fusion to generate a segmentation mask picture with a confidence score; The SAM image segmentation model outputs the position information of each escalator related area relative to the corresponding monitoring scene picture based on the segmentation mask picture with the confidence score.
[0008] Further, the generation process of the segmentation mask picture is:
[0009] Among them, is the finally generated segmentation mask picture, which is used to determine the area of the escalator in the monitoring scene picture; For the input image, i.e., the monitoring scene picture; For the prompt word set; For the parameters of the model.
[0010] Further, the multi-modal large model based on the thought chain is used for prompt word self-adaptive matching of the monitoring scene picture, and core detection is performed according to the generated segmentation mask picture and the matched prompt word to determine whether a high-risk abnormal behavior event such as a person falling down occurs in the escalator area in the station, including: The multi-modal large model based on the thought chain is used to analyze the scene and resolution features of the picture to obtain resolution classification results and escalator quantity classification results of the picture; A dynamic self-adaptive Prompt selection model is constructed, and the corresponding prompt word is self-adaptively matched based on the resolution classification results and the escalator quantity classification results of the monitoring scene picture; The generated segmentation mask picture and the matched prompt word are input into the multi-modal large model based on the thought chain for core detection to determine whether a high-risk abnormal behavior event such as a person falling down occurs in the escalator area in the station.
[0011] Further, the multi-modal large model based on the thought chain is used to analyze the scene and resolution features of the picture to obtain resolution classification results and escalator quantity classification results of the picture, including: The multi-modal large model based on the thought chain is used to classify the monitoring scene picture according to the resolution, and the monitoring scene picture is divided into high-definition pictures and standard-definition pictures to obtain resolution classification results of the monitoring scene picture, wherein the high-definition pictures include high-definition single-escalator pictures and high-definition double-escalator pictures, and the standard-definition pictures include standard-definition single-escalator pictures and standard-definition double-escalator pictures; The number of escalators in the monitoring scene picture is analyzed, and the monitoring scene picture is divided into multi-escalator pictures and single-escalator pictures to obtain escalator quantity classification results of the monitoring scene picture; The resolution classification results and the escalator quantity classification results of the monitoring scene picture are organically integrated to form a composite classification selection basis as a reference dimension in the classification task of the large model.
[0012] Further, the dynamic self-adaptive Prompt selection model is:
[0013] Among them, is the prompt word with the highest confidence in the prompt word set, and are representations of scene features and resolution features, respectively; is the embedding representation of the prompt word set; is a similarity calculation function; and a weight parameter balancing the influence of the scene feature and the resolution feature.
[0014] Further, the generated segmentation mask picture and the matched prompt word are input into a multi-modal large model based on a thinking chain for core detection to determine whether a high-risk abnormal behavior event such as a person falling occurs in the escalator area in the station. The generated segmentation mask picture is input into an image encoder of the multi-modal large model based on the thinking chain to obtain an encoded picture vector. The matched prompt word is input into a text encoder of the multi-modal large model based on the thinking chain to obtain an encoded text vector. The obtained picture vector and text vector are subjected to cross-modal feature alignment to map the visual feature to a language model space to obtain a fused feature vector. The obtained fused feature vector is input into a large language model decoder for inference output to determine whether a high-risk abnormal behavior event occurs in the escalator area in the station, and an alarm is given when a high-risk abnormal behavior event occurs in the escalator in the station.
[0015] In a second aspect, a safety monitoring system for an escalator in a station is provided, comprising: A segmentation mask generation module is configured to obtain a monitoring video of an escalator in the station, and generate a segmentation mask picture based on an input prompt word set and a mask feature map by using a SAM image segmentation model, so as to determine an escalator-related area corresponding to each monitoring scene picture in the monitoring video and obtain position information of each escalator-related area relative to the corresponding monitoring scene picture. A core detection module is configured to use a multi-modal large model based on a thinking chain to adaptively match a prompt word for a monitoring scene picture, and perform core detection based on the generated segmentation mask picture and the matched prompt word to determine whether a high-risk abnormal behavior event such as a person falling occurs in the escalator area in the station. An alarm position determination module is configured to determine an alarm position and an alarm level of the escalator in the station based on an alarm result of the multi-modal large model based on the thinking chain and the position information of each escalator-related area relative to the corresponding monitoring scene picture when a high-risk abnormal behavior event occurs in the escalator area in the station.
[0016] In a third aspect, a processing device is provided, comprising computer program instructions, wherein the computer program instructions are used to implement steps corresponding to the safety monitoring method for an escalator in a station when executed by the processing device.
[0017] In a fourth aspect, a computer readable storage medium is provided, wherein the computer readable storage medium stores computer program instructions, and the computer program instructions are used to implement steps corresponding to the safety monitoring method for an escalator in a station when executed by a processor.
[0018] The present application has the following advantages due to the above technical solutions: 1、The present application adopts a pre-processing mechanism, embeds an image segmentation algorithm in front of a multi-modal large model, performs accurate regional division on the monitoring picture through the algorithm, pre-identifies the spatial boundary of the escalator and the stairs, and thus only inputs the visual information of the escalator region into the large model for analysis. This hierarchical processing architecture effectively makes up for the insufficient sensitivity of the large model to spatial position information, and significantly improves the regional positioning accuracy and overall reliability of abnormal event detection.
[0019] 2、The present application introduces a thought chain (Chain-of-Thought) and constructs a dynamic interactive detection framework. The large model first deeply analyzes the scene features and resolution in the video stream, accurately identifies scene parameters such as the number of escalators and layout structure, and quantifies the image clarity and information density simultaneously. Based on the above analysis results, the present application dynamically matches the optimal Prompt template through a reasoning decision mechanism, realizes accurate adaptation of the detection strategy to the scene features and image quality, and effectively improves the detection robustness and task adaptability of the model under multi-source heterogeneous data.
[0020] 3、The present application can adaptively select corresponding prompt words for different escalator scenes, solving the problem that a single prompt word cannot adapt to multiple scenes, and having strong generalization.
[0021] 4、The present application can not only be used for escalator scenes, but also be applied to other monitoring fields, and has high universality.
[0022] 5、The present application can realize real-time monitoring and analysis of the escalator region, effectively alleviate the artificial pressure, and has high intelligent level.
[0023] In summary, the present application can be widely applied to railway stations, subway stations, shopping malls, airports and other public places to optimize escalator safety monitoring. BRIEF DESCRIPTION OF DRAWINGS
[0024] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments and are not intended to limit the scope of the present application. Throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 is a method flowchart provided by an embodiment of the present application; Figure 2 is a schematic diagram of an image segmentation algorithm provided by an embodiment of the present application; Figure 3 is a scene and resolution feature analysis schematic diagram provided by an embodiment of the present application; Figure 4 is a schematic diagram of prompt adaptive matching provided by an embodiment of the present application; Figure 5 is a schematic diagram of large model detection provided by an embodiment of the present application. DETAILED DESCRIPTION
[0025] Exemplary embodiments of the present application will be described more fully hereinafter with reference to the accompanying drawings, in which exemplary embodiments of the application are shown. This application may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the application to those skilled in the art.
[0026] It should be understood that the terms used herein are for the purpose of describing particular example embodiments and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises" and / or "comprising," and / or "includes" and / or "including" and / or "has" and / or "having" when used herein, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order in which they are described, unless specifically identified as an order dependent step. It is also to be understood that additional or alternative steps can be employed.
[0027] Although the terms first, second, third, and the like can be used herein to describe various elements, components, regions, layers and / or sections, these elements, components, regions, layers and / or sections should not be limited by these terms. These terms can be only used to distinguish one element, component, region, layer or section from another region, layer or section. Terms such as "first," "second," and other numerical terms when used herein do not imply a sequence or order unless clearly indicated by the context. Thus, a first element, component, region, layer or section discussed below could be termed a second element, component, region, layer or section without departing from the teachings of the example embodiments.
[0028] Currently, the multimodal large model has significant limitations in the abnormal event detection task, and its perception ability of spatial position information is weak, which cannot effectively distinguish events occurring in different areas. In the escalator safety monitoring scene, the monitoring picture usually contains both the escalator and the staircase area. When an abnormal situation such as falling occurs in the staircase area, the multimodal large model will misjudge it as an escalator area event, causing the alarm system to be frequently mis-triggered, which seriously affects the accuracy and reliability of the detection. In the escalator monitoring scene, due to the difference in camera types, the resolution of the collected pictures presents diversity. At the same time, due to the existence of different scene configurations such as single escalator operation and multi-escalator operation in the escalator area, the traditional single Prompt (prompt word, instruction of large model) is difficult to adapt to the complex and changeable detection requirements, which seriously restricts the generalization ability and accuracy of abnormal event detection. The embodiment of the present application provides a station escalator safety monitoring method, comprising: acquiring monitoring video of a station escalator, and using a SAM image segmentation model, based on an input prompt word set and a mask feature map, generating a segmentation mask picture to determine the corresponding escalator related area of each monitoring scene picture in the monitoring video, and obtaining the position information of each escalator related area relative to the corresponding monitoring scene picture; using a multimodal large model based on a thinking chain, performing prompt word adaptive matching on the monitoring scene picture, and performing core detection according to the generated segmentation mask picture and the matched prompt word to determine whether a high-risk abnormal behavior event such as falling of personnel occurs in the station escalator area; when a high-risk abnormal behavior event occurs in the station escalator area, determining the alarm position and alarm level of the station escalator according to the alarm result of the multimodal large model based on the thinking chain and the position information of each escalator related area relative to the corresponding monitoring scene picture. The present application can significantly improve the regional positioning accuracy and overall reliability of abnormal event detection.
[0029] Embodiment 1 As shown in Figure 1 , the embodiment provides a station escalator safety monitoring method, comprising the following steps: 1) acquiring monitoring video of a station escalator, and using a SAM image segmentation model, based on an input prompt word set and a mask feature map, generating a segmentation mask picture to determine the corresponding escalator related area of each monitoring scene picture in the monitoring video, and obtaining the position information of each escalator related area relative to the corresponding monitoring scene picture, specifically: 1.1) acquiring monitoring video of a station escalator.
[0030] 1.2) performing frame extraction operation on the acquired monitoring video to obtain the corresponding monitoring scene picture.
[0031] 1.3) input each monitoring scene picture into an image encoder of a SAM (Segment Anything Model) image segmentation model for processing, and convert each monitoring scene picture into an image embedding.
[0032] Specifically, the SAM image segmentation model is based on a Transformer architecture, as shown in the following formula (1) : Figure 2 The monitoring scene picture is input into the image encoder of the SAM image segmentation model to convert the monitoring scene picture into an image embedding, so as to realize accurate understanding and efficient processing of image content. The process of converting the monitoring scene picture into the image embedding can be represented as:
[0033] wherein, is the obtained image embedding; is the input image, i.e., the monitoring scene picture, is a model parameter.
[0034] 1.4) input the text prompt word (for example, escalator in the present application, which is used to guide the SAM image segmentation model to segment the image region that the model wants) into a Prompt encoder of the SAM image segmentation model for processing, and convert the text prompt word into a prompt embedding.
[0035] 1.5) input the mask feature map into a down-sampling module of the SAM image segmentation model for processing, to obtain a down-sampled mask feature map.
[0036] Specifically, the down-sampling module progressively reduces the dimension through two 2x2 convolution layers with a step of 2, and maps the channels through a 1x1 convolution. The purpose of down-sampling the mask feature map is that the feature map output by the image encoder is 1 / 16 of the original monitoring scene picture, so the mask also needs to be reduced accordingly to correspond to the image embedding.
[0037] 1.6) input the converted image embedding and prompt embedding and the down-sampled mask feature map into a lightweight mask decoder of the SAM image segmentation model for efficient fusion, to generate a segmentation mask picture with a confidence score, and realize accurate identification of the escalator-related region in each monitoring scene picture.
[0038] Specifically, the down-sampled mask feature map is input into the lightweight mask decoder together with the transformed image embedding and the prompt embedding. The lightweight mask decoder integrates the outputs of the image encoder and the prompt encoder through the self-attention mechanism and the cross-attention mechanism of the Transformer, and deepens the down-sampled mask feature map into the reasoning process. The lightweight mask decoder fuses these features, so that the SAM image segmentation model can comprehensively consider the image content, spatial position information, and position and shape information of the mask prompt, thereby more accurately generating a segmentation mask picture of the target object. In this process, the down-sampled mask feature map provides prior information of the mask and a rough segmentation region, and the lightweight mask decoder generates a more accurate segmentation mask picture through step-by-step upsampling and refinement.
[0039] Specifically, the generation process of the segmentation mask picture is as follows:
[0040] Wherein, is the final generated segmentation mask picture, which is used to determine the accurate region of the escalator in the monitoring scene picture; is a set of prompt words (including a plurality of pre-set prompt words, such as boundary points of the escalator); is a parameter of the model.
[0041] 1.7) Based on the segmentation mask picture with the confidence score, the SAM image segmentation model appropriately enlarges the area, fully ensures the complete coverage of the escalator scene, and outputs the position information of each escalator related region relative to the corresponding monitoring scene picture, thereby providing key regional basic information for subsequent passenger behavior detection or abnormal situation identification tasks.
[0042] 2) A multi-modal large model based on a thinking chain is adopted to adaptively match the prompt words (Prompt) of the monitoring scene picture, and a core detection is performed according to the generated segmentation mask picture and the matched prompt words to determine whether a high-risk abnormal behavior event such as falling occurs in the escalator region in the station.
[0043] Specifically, the core principle of the thinking chain introduced by the present application is to simulate the thinking process of humans when solving problems, and to decompose complex problems into a series of intermediate reasoning steps. In terms of implementation, a prompt engineering is usually adopted, that is, an example or a guiding sentence is provided for the model in the input. Therefore, the specific process of this step is as follows: 2.1) As shown in Figure 3 , a multi-modal large model based on a thinking chain is adopted to analyze the scene and resolution characteristics of the picture, and obtain the resolution classification result and the escalator number classification result of the picture: 2.1.1) Adopting a multi-modal large model based on thought chain, classify the monitoring scene pictures according to resolution, divide the monitoring scene pictures into high-definition pictures and standard-definition pictures, and obtain the resolution classification result of the monitoring scene pictures, wherein the high-definition pictures include high-definition single escalator pictures and high-definition double escalator pictures, and the standard-definition pictures include standard-definition single escalator pictures and standard-definition double escalator pictures.
[0044] Specifically, resolution as a key indicator to measure the quality of pictures, can reflect the details and information richness included in the pictures, high-definition pictures can often provide more rich textures and more delicate image features, while standard-definition pictures are relatively low in data volume and clarity.
[0045] Specifically, when classifying the monitoring scene pictures according to resolution, the pictures are classified according to the relationship between the picture size of the monitoring scene pictures and the resolution threshold value:
[0046] Among them, is the classification result according to resolution (1 for high-definition, 0 for standard-definition).
[0047] 2.1.2) Analyze the number of escalators in the monitoring scene pictures, divide the monitoring scene pictures into multi-escalator pictures and single-escalator pictures, and obtain the escalator number classification result of the monitoring scene pictures.
[0048] Specifically, the different number of escalators means that the complexity of the scene is different, the single-escalator scene is relatively simple, while the multi-escalator scene involves more complex factors such as the spatial relationship between escalators, layout, etc.
[0049] Specifically, when classifying the monitoring scene pictures according to the number of escalators, the pictures are classified according to the detected number of escalators and the number threshold value:
[0050] Among them, is the classification result according to the number of escalators.
[0051] 2.1.3) Organically integrate the resolution classification result and the escalator number classification result of the monitoring scene pictures to form a kind of composite classification selection basis as a reference dimension in the large model classification task.
[0052] 2.2) As shown in Figure 4 , a dynamic adaptive Prompt (prompt word) selection model is constructed, and based on the resolution classification result and the escalator number classification result of the monitoring scene pictures, the corresponding prompt word is adaptively matched: 2.2.1) Constructing a dynamic adaptive Prompt selection model:
[0053] wherein, is the prompt with the highest confidence in the prompt set, and are the representations of the scene feature and resolution feature, respectively; is the embedding representation of the prompt set ; is a similarity calculation function (such as cosine similarity); and are weight parameters that balance the influence of scene features and resolution features. This formula is used to select the prompt that best matches the current scene features and resolution features from the prompt set, to achieve dynamic adaptive Prompt selection.
[0054] 2.2.2) Using the constructed dynamic adaptive Prompt selection model, based on the resolution classification results and escalator quantity classification results of the monitoring scene pictures, to adaptively match the prompt corresponding to the monitoring scene pictures.
[0055] Specifically, after the multi-modal large model based on thought chain analyzes the scene and resolution features of the monitoring scene pictures, it can accurately determine the category to which it belongs, and then retrieve the matching prompt. The core logic of adaptive Prompt is that the lower the image information density and the more complex the scene, the more the prompt needs to compensate for information loss through structured constraints.
[0056] Specifically, the dynamic adaptive Prompt selection model includes first, second, third, and fourth levels, with the first level being the most lenient, the second level being relatively lenient, the third level being relatively strict, and the fourth level being the most strict: For high-definition single-escalator pictures, the designed prompt is the first level, aiming to fully exploit the potential information in the image: the prompt has no format restrictions, allowing the model to exercise its autonomy, and only needs to provide the most basic task description, such as "analyze the passenger behavior in the escalator area and determine whether there is an abnormal state."
[0057] For high-definition double-escalator pictures, although the image quality is high, it contains multiple escalators, and therefore needs to balance information capture and accuracy, so the designed prompt is the second level: the prompt needs to add some scene information description in addition to the most basic task description, guiding the model to focus on scene information, such as "there are multiple escalator areas in the picture, please analyze whether there are abnormal passenger behaviors in these escalator areas."
[0058] For standard definition escalator images, due to their lower image resolution, more precise positioning and identification are required. Therefore, the designed prompts are at the third level: in addition to the task description and scene information mentioned above, the prompts add some constraints to avoid false alarms caused by the lack of key information in standard definition images, such as "The image is a standard definition image with low resolution. Please analyze whether there is any abnormal passenger behavior in the escalator area of the image. If you are unsure, please do not answer."
[0059] For standard definition (SD) images of two escalators, the detection is most difficult due to their low resolution and numerous elements. Therefore, the prompts are designed at level four: the prompts have the most restrictions and require a structured descriptive framework, such as: "1. Prompt: The image is SD with low resolution and contains multiple escalator areas; 2. Task: For each escalator area, determine whether there is any abnormal behavior by personnel; 3. Judgment criteria: Judgment is based solely on clearly identifiable visual features, without inferring the behavior of blurred areas; 4. Output: Output "Yes" if there is abnormal behavior, and "No" if there is no abnormal behavior; 5. Prohibition: Do not make unfounded speculations, do not describe non-abnormal behaviors (such as normal standing or walking), and do not output conclusive statements for uncertain content."
[0060] 2.3) As Figure 5 As shown, the generated segmentation mask image and the matched prompt are input into a multimodal large model based on thought chain for core detection to determine whether high-risk abnormal behavior events such as people falling have occurred in the escalator area of the station. 2.3.1) Input the generated segmentation mask image into the image encoder based on the multimodal large model of the thought chain to obtain the encoded image vector.
[0061] 2.3.2) Input the matched prompt words into the text encoder of the multimodal large model based on the thought chain to obtain the encoded text vector.
[0062] 2.3.3) Perform cross-modal feature alignment on the obtained image vectors and text vectors, mapping the visual features to the language model space to obtain the fused feature vectors.
[0063] 2.3.4) Input the obtained fused feature vector into the large language model decoder for inference output to determine the detection result, that is, to determine whether a high-risk abnormal behavior event such as a person falling has occurred in the escalator area of the station. When a high-risk abnormal behavior event occurs in the escalator of the station, an alarm is pushed.
[0064] Specifically, the core task of the thought chain-based multi-modal large model at this stage is to perform intelligent monitoring and analysis of the running state of the station escalator, focusing on detecting and identifying whether there is a high-risk abnormal behavior event such as falling down. Based on the input image and the fine-guided prompt word, the thought chain-based multi-modal large model will output the analysis result to provide key basis for safety monitoring decision.
[0065] 3) When a high-risk abnormal behavior event occurs in the station escalator area, according to the alarm result of the thought chain-based multi-modal large model and the position information of each escalator related area relative to the corresponding monitoring scene picture, the alarm position and alarm level of the station escalator are determined, so that the on-site personnel can quickly respond and avoid further expansion of the danger.
[0066] Specifically, the alarm level responds differently according to different high-risk abnormal behavior events, including first-level alarm and second-level alarm. The high-risk abnormal behavior event of the first-level alarm is to carry large luggage, baby strollers, etc. on the escalator; the high-risk abnormal behavior event of the second-level alarm is to fall down in the escalator related area. The first-level alarm only needs to notify the monitoring personnel to pay attention to the high-risk event in the escalator related area; the second-level alarm notifies the monitoring personnel, and at the same time, since the event has occurred, the start and stop of the escalator will be controlled through the remote button to avoid the occurrence of danger in time.
[0067] Embodiment 2 The embodiment provides a station escalator safety monitoring system, comprising: The segmentation mask generation module is configured to obtain a monitoring video of a station escalator, and generate a segmentation mask picture based on an input prompt word set and a mask feature map by using a SAM image segmentation model, so as to determine corresponding escalator related areas of each monitoring scene picture in the monitoring video, and obtain position information of each escalator related area relative to the corresponding monitoring scene picture.
[0068] The core detection module is configured to use a thought chain-based multi-modal large model to perform prompt word self-adaptive matching on the monitoring scene picture, and perform core detection according to the generated segmentation mask picture and the matched prompt word, so as to determine whether a high-risk abnormal behavior event such as falling down occurs in the station escalator area.
[0069] The alarm position determination module is configured to, when a high-risk abnormal behavior event occurs in the station escalator area, determine the alarm position and alarm level of the station escalator according to the alarm result of the thought chain-based multi-modal large model and the position information of each escalator related area relative to the corresponding monitoring scene picture.
[0070] The system provided in the embodiment is used to execute the above-mentioned method embodiments, and the specific process and detailed content are referred to the above-mentioned embodiments, which will not be described here again.
[0071] Embodiment 3 This embodiment provides a processing device corresponding to the escalator safety monitoring method provided in Embodiment 1, which can be applied to a client processing device, such as a mobile phone, a notebook computer, a tablet computer, a desktop computer, etc., to execute the method of Embodiment 1.
[0072] The processing device includes a processor, a memory, a communication interface, and a bus, and the processor, the memory, and the communication interface are connected through the bus to complete communication with each other. The memory stores a computer program that can run on the processing device, and the processing device executes the escalator safety monitoring method provided in Embodiment 1 when running the computer program.
[0073] In some implementations, the memory can be a high-speed random access memory (RAM), and can also include a non-volatile memory, such as at least one disk memory.
[0074] In other implementations, the processor can be a central processing unit (CPU), a digital signal processor (DSP), or various types of general-purpose processors, which are not limited here.
[0075] In addition, the logical instructions in the memory described above can be implemented in the form of a software functional unit and sold or used as a separate product, which can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0076] Those skilled in the art can understand that the structure of the computing device described above is only part of the structure related to the present application scheme, and does not constitute a limitation on the computing device to which the present application scheme is applied. The specific computing device can include more or fewer components, or combine certain components, or have a different component arrangement.
[0077] Embodiment 4 The embodiment provides a computer program product corresponding to the escalator safety monitoring method in the station provided in the embodiment 1, and the computer program product can include a computer readable storage medium, which is loaded with computer readable program instructions for executing the escalator safety monitoring method in the station described in the embodiment 1.
[0078] The computer readable storage medium can be a tangible device that maintains and stores instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.
[0079] The computer readable storage medium provided in the above embodiment has similar implementation principles and technical effects to the method embodiments, and details are not described herein.
[0080] The present application is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one block or multiple blocks.
[0081] These computer program instructions can also be stored in a computer readable memory that can guide the computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one block or multiple blocks.
[0082] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one block or multiple blocks.
[0083] The above embodiments are only used for illustrating the present application, wherein the structure, connection mode and manufacturing process of each component can be changed, and any equivalent transformation and improvement based on the technical scheme of the present application should not be excluded from the protection scope of the present application.
Claims
1. A method of escalator safety monitoring in a station, characterized by The application relates to a method for monitoring high-risk abnormal behavior events of an escalator in a station, and a device thereof. The method comprises the following steps: acquiring monitoring videos of the escalator in the station, and adopting a SAM image segmentation model to generate segmentation mask pictures based on an input set of prompt words and a mask feature map, so as to determine the escalator-related regions corresponding to the monitoring scene pictures in the monitoring videos and obtain the position information of the escalator-related regions relative to the corresponding monitoring scene pictures; adopting a multi-modal large model based on a thinking chain to adaptively match the prompt words for the monitoring scene pictures, and performing core detection according to the generated segmentation mask pictures and the matched prompt words, so as to determine whether high-risk abnormal behavior events occur in the escalator region in the station; and when high-risk abnormal behavior events occur in the escalator region in the station, determining the alarm position and alarm level of the escalator in the station according to the alarm result of the multi-modal large model based on the thinking chain and the position information of the escalator-related regions relative to the corresponding monitoring scene pictures. The method for acquiring the monitoring videos of the escalator in the station and generating the segmentation mask pictures based on the input set of prompt words and the mask feature map so as to determine the escalator-related regions corresponding to the monitoring scene pictures in the monitoring videos and obtain the position information of the escalator-related regions relative to the corresponding monitoring scene pictures comprises the following steps: acquiring the monitoring videos of the escalator in the station; 2. A method of monitoring the safety of a moving stairway in a station as claimed in claim 1, characterized in that, performing frame extraction on the acquired monitoring videos to obtain corresponding monitoring scene pictures; inputting each monitoring scene picture into an image encoder of a SAM image segmentation model for processing, so as to convert each monitoring scene picture into an image embedding; inputting a text prompt word into a Prompt encoder of the SAM image segmentation model for processing, so as to convert the text prompt word into a prompt embedding; inputting a mask feature map into a down-sampling module of the SAM image segmentation model for processing to obtain a down-sampled mask feature map; inputting the converted image embedding and prompt embedding and the down-sampled mask feature map into a lightweight mask decoder of the SAM image segmentation model for fusion to generate a segmentation mask picture with a confidence score; the SAM image segmentation model outputs the position information of each escalator-related region relative to the corresponding monitoring scene picture based on the segmentation mask picture with the confidence score. The generation process of the segmentation mask picture comprises the following steps: the method for adopting the multi-modal large model based on the thinking chain to adaptively match the prompt words for the monitoring scene pictures and performing core detection according to the generated segmentation mask picture and the matched prompt words so as to determine whether high-risk abnormal behavior events such as falling occur in the escalator region in the station comprises the following steps:
3. A method of monitoring the safety of a station escalator as claimed in claim 2, wherein, adopting the multi-modal large model based on the thinking chain to analyze the scene and resolution features of a picture to obtain a resolution classification result and an escalator quantity classification result of the picture; ; wherein, is the final generated segmentation mask picture, used to determine the area of the escalator in the monitoring scene picture; is the monitoring scene picture; is the prompt word set; is the parameter of the model.
4. A method of monitoring the safety of a moving stairway in a station as defined in claim 1, characterized by constructing a dynamic adaptive Prompt selection model, and adaptively matching corresponding prompt words based on the resolution classification result and the escalator quantity classification result of the monitoring scene picture; inputting the generated segmentation mask picture and the matched prompt words into the multi-modal large model based on the thinking chain for core detection, so as to determine whether high-risk abnormal behavior events such as falling occur in the escalator region in the station. 5. A method of monitoring the safety of a moving stairway in a station as claimed in claim 4, characterized in that, The multi-modal large model based on the thought chain is used to analyze the scene and resolution characteristics of the picture, and resolution classification results and escalator quantity classification results of the picture are obtained, including: The multi-modal large model based on the thought chain is used to classify the monitoring scene pictures according to the resolution, and the monitoring scene pictures are divided into high-definition pictures and standard-definition pictures, and resolution classification results of the monitoring scene pictures are obtained, wherein the high-definition pictures include high-definition single-escalator pictures and high-definition double-escalator pictures, and the standard-definition pictures include standard-definition single-escalator pictures and standard-definition double-escalator pictures; The number of escalators in the monitoring scene pictures is analyzed, and the monitoring scene pictures are divided into multi-escalator pictures and single-escalator pictures, and escalator quantity classification results of the monitoring scene pictures are obtained; The resolution classification results and the escalator quantity classification results of the monitoring scene pictures are organically integrated to form a composite classification selection basis as a reference dimension in the large model classification task.
6. A method of monitoring the safety of a moving stairway in a station as defined in claim 4, characterized by The dynamic self-adaptive prompt selection model is: ; wherein, is a prompt word in the prompt word set, and are representations of the scene feature and the resolution feature, respectively; is an embedding representation of the prompt word set; is a similarity calculation function; and are weight parameters balancing the influence of the scene feature and the resolution feature.
7. A method of monitoring the safety of a moving stairway in a station as defined in claim 4, characterized by The generated segmentation mask picture and the matched prompt word are input into the multi-modal large model based on the thought chain for core detection to determine whether a high-risk abnormal behavior event occurs in the escalator area in the station, including: The generated segmentation mask picture is input into the image encoder of the multi-modal large model based on the thought chain to obtain an encoded picture vector; The matched prompt word is input into the text encoder of the multi-modal large model based on the thought chain to obtain an encoded text vector; The obtained picture vector and text vector are subjected to cross-modal feature alignment to map the visual features to the language model space to obtain a fused feature vector; The obtained fused feature vector is input into the large language model decoder for inference output to determine whether a high-risk abnormal behavior event occurs in the escalator area in the station, and an alarm is given when a high-risk abnormal behavior event occurs in the escalator in the station.
8. An escalator safety monitoring system in a station, characterized by It includes: A segmentation mask generation module is configured to obtain a monitoring video of an escalator in a station, and generate a segmentation mask picture based on an input prompt word set and a mask feature map using a SAM image segmentation model, to determine the escalator-related area corresponding to each monitoring scene picture in the monitoring video and obtain the position information of each escalator-related area relative to the corresponding monitoring scene picture; A core detection module is configured to use a multi-modal large model based on a thought chain to adaptively match the prompt words for the monitoring scene pictures, and to perform core detection based on the generated segmentation mask picture and the matched prompt words to determine whether a high-risk abnormal behavior event occurs in the escalator area in the station; An alarm position determination module is configured to determine the alarm position and alarm level of the escalator in the station based on the alarm result of the multi-modal large model based on the thought chain and the position information of each escalator-related area relative to the corresponding monitoring scene picture when a high-risk abnormal behavior event occurs in the escalator area in the station.
9. A processing device, characterized by It includes computer program instructions, wherein the computer program instructions are executed by a processing device to implement the steps corresponding to the escalator safety monitoring method in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions are used for realizing the steps corresponding to the escalator safety monitoring method in any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Driving risk early warning method and system based on large language model
CN118701092A
Escalator monitoring and early warning method and device based on target motion change
CN119520945A
Civil engineering structure apparent damage diagnosis method based on multi-modal large model
CN119785098A
Panoramic vision relation detection method based on thinking chain reasoning
CN120655919A
Method, apparatus, device and medium for multimodal data processing
US20250182286A1