Elevator stuck risk level assessment method based on multi-modal large model
By automatically evaluating the risk level of entrapment when an elevator stops through a system based on a multimodal large model, the problem of the existing technology that cannot effectively distinguish the entrapment risks of different groups of people is solved, and efficient and safe elevator fault handling is achieved.
Patent Information
- Application Number
- CN202411982836.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing technologies are unable to effectively distinguish and assess the risk levels of entrapment for different groups of people when an elevator stops, resulting in inefficiency for customer service personnel in handling entrapment incidents.
A system based on a multimodal large model is used to automatically collect car images through the DTU terminal, communication server and image analysis module. The multimodal large model is used to describe the images and convert them into natural language text. The risk level is evaluated and the processing priority is determined in combination with the risk assessment model, and a fault handling order is automatically issued.
It improves the efficiency of handling trapped people faults, reduces the workload of customer service personnel, ensures the safety of elevator riding, and can promptly confirm whether there are people trapped in the car and determine the handling priority.
Smart Images

Figure CN119911768B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of elevator safety monitoring, and in particular relates to a method for assessing the risk level of people trapped in an elevator due to a stoppage based on a multimodal large model. Background Art
[0002] With the development of cities and the improvement of living standards, elevators are used more and more frequently in our lives. As a vertical transportation tool, elevators may malfunction and stop for various reasons during use. When the elevator stops, there may be passengers in the elevator car, which may be trapped. Passengers trapped in the car are in danger (psychologically and physiologically) to a certain extent and must be rescued in time to ensure their safety. In actual work, when an elevator stops, customer service personnel or sensor data will determine whether there are people trapped. If there are people trapped, they must be handled as a priority to ensure their timely rescue. However, the risk level faced by normal adults and special groups (elderly, children, disabled people, etc.) is different.
[0003] Chinese invention patent CN118270617A discloses a method and system for early warning of elevator operation and entrapment risk based on deep learning, comprising the following steps: Step S101, constructing a feature matrix; Step S102, constructing a feature sequence. Step S103, inputting the feature matrix and feature sequence into the elevator risk assessment model, and the output value represents the elevator risk score within the preset prediction time period K; Step S104, taking the first safety early warning measure according to the elevator risk score; the system includes a feature matrix construction module, which is used to set M image sampling points in the elevator car, machine room, and shaft, and collect the first image data of the elevator at a preset time interval n within the preset sampling time period m, and construct a feature matrix; a feature sequence construction module, which is used to collect the operation data and second image data of the elevator at a preset time interval n within the preset collection time period m, and construct a feature sequence; an elevator risk score output module, which is used to input the feature matrix and feature sequence into the elevator risk assessment model, and the output value represents the elevator risk score within the preset prediction time period K. Risk score; a first safety warning measure generation module, which is used to take the first safety warning measure according to the elevator risk score; a manned image sequence generation module, which is used to extract the image inside the elevator car from the second image data at each time point in the second sequence to generate a manned image sequence; a maximum angle value calculation module, which is used to input the manned image sequence into the manned risk assessment model, and the output value represents the maximum angle value between the torso key point of the person in the manned image sequence and the elevator plane; a second safety warning measure generation module, which is used to take the second safety warning measure according to the maximum angle value; this solution uses a neural network model to extract features from the elevator image, and determines whether there is a dangerous situation through the angle of the limb key points, but this solution does not grade the risk level of entrapment, which is not convenient for subsequent customer service personnel to handle entrapment incidents.
[0004] Therefore, the present invention provides a method for assessing the risk level of people trapped in an elevator due to an elevator stall based on a multimodal large model to solve the problems raised by the above background technology. Summary of the Invention
[0005] In response to the problems raised by the above-mentioned background technology, the purpose of the present invention is to provide a method for assessing the risk level of people trapped in an elevator due to a stop-and-entrapment based on a multimodal large model. After an elevator stops, images inside the car are automatically obtained, and the risk level of people trapped is assessed through the multimodal large model and the risk assessment model. Then, the processing priority is determined, thereby reducing the workload of customer service personnel, improving the efficiency of handling people trapped faults, and ensuring the safety of elevator riding.
[0006] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:
[0007] A system for assessing the risk level of trapped people in an elevator stall based on a multimodal large model, the system comprising a DTU terminal, a communication server, an image analysis module, and an automatic ordering module, wherein the DTU terminal is communicatively connected to the communication server, the communication server is communicatively connected to the image analysis module, and the image analysis module is communicatively connected to the automatic ordering module;
[0008] The DTU terminal is used to collect elevator operation status data and fault data in real time and send them to the communication server in real time. At the same time, the DTU terminal can perform image acquisition according to the instructions of the communication server;
[0009] The communication server is used to receive the elevator operation status data and fault data uploaded by the DTU terminal, and control the DTU terminal to collect the car image when receiving the elevator stop fault data;
[0010] The image analysis module is used to evaluate the risk level of the elevator car image and determine the processing priority;
[0011] The automatic ordering module is used to automatically send the fault handling order to the maintenance personnel.
[0012] The risk level assessment method for entrapment in elevator stalls based on a multimodal large model includes the following steps:
[0013] S1: receiving data uploaded by the elevator through the communication server, and capturing images when elevator stop fault data is received;
[0014] S2: The communication server sends a command to collect car images, and the DTU controls the camera in the car to take pictures and transmit them back to the communication server;
[0015] S3: After receiving the photo of the cabin, the communication server sends it to the image analysis module, and the image analysis module uses the multimodal large model to obtain a description of the image;
[0016] S4: The image analysis module obtains the image description and sends it to the risk assessment model for risk assessment to obtain the risk level;
[0017] S5: The image analysis module determines the processing priority according to the risk level, and the higher the risk, the higher the priority;
[0018] S6: The automatic ordering module determines whether to automatically place an order based on the configuration.
[0019] It is further defined that when the communication server in S1 receives the elevator stop fault data, it first checks whether the DTU terminal supports the image acquisition function. If it does, it starts to acquire images. If it does not support image acquisition, the customer service staff confirms the elevator stop situation.
[0020] It is further defined that the risk levels of the risk assessment model in S4 include 0 to 3 levels, where 0 indicates no risk, there is no one and no items in the elevator car; 1 indicates low risk, there is no one but there are items in the elevator car; 2 indicates medium risk, there is an adult male in the elevator car; 3 indicates high risk, there are elderly people, children, women or disabled people in the elevator car.
[0021] It is further defined that the multimodal large model in S3 is used to describe the collected images inside the car, thereby obtaining information inside the car described using natural language, and the natural language is input into the risk assessment model.
[0022] It is further defined that after receiving the natural language description text output by the multimodal large model, the risk assessment model performs word segmentation on the text, that is, dividing a sentence into multiple word sequences, and sending the word sequences into the risk assessment model in sequence.
[0023] It is further defined that the model training process of the risk assessment model in S4 includes the following steps:
[0024] S41: Label the pictures, label the collected car pictures, and mark the risk level according to the situation of the car pictures;
[0025] S42: multimodal large model description, calling the multimodal large model to describe each car image and save the corresponding description text;
[0026] S43: Integrate the description text and the label, and integrate the description text and risk level of the image to form a training sample.
[0027] It is further defined that the model reasoning process of the risk assessment model in S4 includes the following steps:
[0028] S411: Receiving a picture, the image analysis module receives a car picture sent by the communication server;
[0029] S412: describing the multimodal large model, calling the multimodal large model to describe the car image and outputting corresponding description text;
[0030] S413: Output the description text to the risk assessment model for risk level assessment.
[0031] It is further defined that if the automatic ordering module in S6 automatically places an order, the ordering process will be automatically called to send the fault handling order to the maintenance personnel. If the order is not automatically placed, the order will be placed after confirmation by the customer service personnel.
[0032] The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model provided by the present invention has the following beneficial effects:
[0033] 1. The application utilizes the picture description capability of a multi-modal large model to realize conversion from pictures to natural language, thereby converting the picture classification problem into a text classification problem, which is more intuitive and simple.
[0034] 2. The application utilizes open source or commercial multi-modal large models, which not only saves the training cost of traditional large models, but also ensures that the description accuracy of the model itself is within an acceptable range, and can also follow the continuous iteration and optimization of the model.
[0035] 3. The application continuously iteratively optimizes the risk assessment model by collecting sample data during actual operation, thereby improving the accuracy and generalization ability of the model.
[0036] 4. The application automatically collects car pictures and performs risk assessment when the elevator fails to stop, can timely confirm whether there is a person trapped in the car and confirm the priority of processing, brings great convenience to customer service personnel, can automatically issue a fault handling sheet to maintenance personnel, reduces the workload of customer service personnel, improves the trapped fault handling efficiency, shortens the fault handling time, ensures the safety of taking the elevator and reduces the loss. BRIEF DESCRIPTION OF DRAWINGS
[0037] The application can be further illustrated by the non-limiting embodiments shown in the accompanying drawings;
[0038] Figure 1 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0039] Figure 2 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0040] Figure 3 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0041] Figure 4 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0042] Figure 5 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0043] Figure 6 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0044] Figure 7 The application is based on the overall structure of the elevator trapped risk level assessment method based on the multi-modal large model embodiment;
[0045] Figure 8 This is an annotated diagram of a text classification model for an embodiment of a method for assessing the risk level of entrapment in an elevator stall based on a multimodal large model according to the present invention;
[0046] Figure 9 This is a diagram of the reasoning process of a text classification model according to an embodiment of the method for assessing the risk level of entrapment in an elevator stall based on a multimodal large model of the present invention;
[0047] Figure 10 This is an example diagram of an elevator car image according to an embodiment of the method for assessing the risk level of entrapment in an elevator due to a stoppage based on a multimodal large model of the present invention. DETAILED DESCRIPTION
[0048] In order to enable those skilled in the art to better understand the present invention, the technical solutions of the present invention are further described below in conjunction with the accompanying drawings and embodiments. The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts are within the scope of protection of the present invention.
[0049] It should be noted that all directional indications in the embodiments of the present invention (such as up, down, left, right, front, back, etc.) are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0050] In addition, the descriptions of "first", "second", etc. in the present invention are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.
[0051] like Figure 1As shown, the system of the risk level assessment method for elevator entrapment based on a multimodal large model of the present invention includes a DTU terminal, a communication server, an image analysis module and an automatic ordering module. The DTU terminal is communicatively connected to the communication server, the communication server is communicatively connected to the image analysis module, and the image analysis module is communicatively connected to the automatic ordering module.
[0052] The DTU terminal is used to collect elevator operation status data and fault data in real time and send them to the communication server in real time. At the same time, the DTU terminal can perform image acquisition according to the instructions of the communication server;
[0053] The communication server is used to receive the elevator operation status data and fault data uploaded by the DTU terminal, and control the DTU terminal to collect the car image when receiving the elevator stop fault data;
[0054] The image analysis module is used to evaluate the risk level of the elevator car image and determine the processing priority;
[0055] The automatic ordering module is used to automatically send the fault handling order to the maintenance personnel.
[0056] like Figure 2 As shown in FIG, a method for assessing the risk level of trapped people in an elevator due to a stopped elevator based on a multimodal large model includes the following steps:
[0057] S1: receiving data uploaded by the elevator through the communication server, and capturing images when elevator stop fault data is received;
[0058] S2: The communication server sends a command to collect car images, and the DTU controls the camera in the car to take pictures and transmit them back to the communication server;
[0059] S3: After receiving the photo of the cabin, the communication server sends it to the image analysis module, and the image analysis module uses the multimodal large model to obtain a description of the image;
[0060] S4: The image analysis module obtains the image description and sends it to the risk assessment model for risk assessment to obtain the risk level;
[0061] S5: The image analysis module determines the processing priority according to the risk level, and the higher the risk, the higher the priority;
[0062] S6: The automatic ordering module determines whether to automatically place an order based on the configuration.
[0063] In the actual application of this embodiment, when the communication server in S1 receives the elevator stop fault data, it first checks whether the DTU terminal supports the image acquisition function. If it does, it starts to acquire images. If it does not support image acquisition, the customer service staff confirms the elevator stop situation.
[0064] In the actual application of this embodiment, the risk level of the risk assessment model in S4 includes levels 0 to 3, where 0 represents no risk, that is, there is no one and no items in the elevator car; 1 represents low risk, that is, there is no one but there are items in the elevator car; 2 represents medium risk, that is, there is an adult male in the elevator car; 3 represents high risk, that is, there are elderly people, children, women or disabled people in the elevator car.
[0065] In the practical application of this embodiment, the multimodal large model in S3 is used to describe the collected images inside the car, thereby obtaining the information inside the car described in natural language, and inputting the natural language into the risk assessment model. Since the training of the multimodal large model requires a large amount of material and computing resources, and the cost-effectiveness of training a multimodal large model specifically for elevator scenarios is very low, a commercial or open source multimodal large model can be used. Figure 10 As shown, the description given by the multimodal large model is: "The person in the photo is a male. He is wearing a black short-sleeved shirt, black shorts, and brown sandals. He is sitting in the elevator car with his hands on his knees, his legs straight, and his body leaning forward. There are yellow and black patterns on the floor of the elevator, which may be safety signs or decorations." It can be seen that the description of the multimodal large model is very accurate, which provides accurate input data for the subsequent risk assessment model.
[0066] In the actual application of this embodiment, the risk assessment model receives the natural language description text output by the multimodal large model and performs word segmentation on the text, that is, dividing a sentence into multiple word sequences, and sending the word sequences into the risk assessment model in sequence.
[0067] like Figure 3 As shown, in the actual application of this embodiment, the risk assessment model in S4 includes 5 layers, the first layer is the embedding layer, and the embedding layer maps each word into a vector; specifically, during the model training process, the embedding vector will be continuously updated, and pre-trained word vectors (such as pre-trained word vectors of Tencent-ailab or Zhihu) can also be used. When using pre-trained word vectors, the embedding layer directly loads the pre-trained word vectors, and the embedding vectors will not be updated during the training process.
[0068] The second layer is a BiLSTM layer, i.e., a bidirectional long short-term memory layer, which is composed of a forward LSTM and a backward LSTM, and is used to better capture bidirectional semantic dependencies.
[0069] The third layer is a normalization layer, which is used to normalize the output data of the second layer (BiLSTM layer); the fourth layer is a ReLU activation function layer, which is used to activate the output data of the normalization layer.
[0070] The fifth layer is a Linear layer, i.e., a linear layer, which is used to perform a fully connected linear transformation on the output data of the activation function layer. The output data of the linear layer is the predicted probability of the risk level category; the risk level assessment is output by the activation function layer using the LogSoftmax function.
[0071] like Figure 4 As shown, in the actual application of this embodiment, the model training process of the risk assessment model in S4 includes the following steps:
[0072] S41: Label the pictures, label the collected car pictures, and mark the risk level according to the situation of the car pictures;
[0073] S42: multimodal large model description, calling the multimodal large model to describe each car image and save the corresponding description text;
[0074] S43: Integrate the description text and the label, and integrate the description text and risk level of the image to form a training sample.
[0075] like Figure 5 As shown, in the actual application of this embodiment, the model reasoning process of the risk assessment model in S4 includes the following steps:
[0076] S411: Receiving a picture, the image analysis module receives a car picture sent by the communication server;
[0077] S412: describing the multimodal large model, calling the multimodal large model to describe the car image and outputting corresponding description text;
[0078] S413: Output the description text to the risk assessment model for risk level assessment.
[0079] In the actual application of this embodiment, if the automatic ordering module in S6 automatically places an order, the ordering process will be automatically called to send the fault handling order to the maintenance personnel. If the order is not automatically placed, the order will be placed after confirmation by the customer service personnel.
[0080] like Figure 6 As shown in the figure, in actual implementation, this system is divided into image analysis services and front-end systems. The front-end system refers to the existing elevator remote monitoring system. The image analysis service includes two processes: model training and deployment and inference. These two processes form a closed loop during the iterative optimization of the model:
[0081] 1. By collecting images of the elevator cabin after the door is closed, these images can be simulated and collected in the elevator. After the images are collected, they are labeled according to the classification criteria of the risk assessment model. Since the risk of human entrapment is divided into 0-3 levels (0: no risk (no one and no objects in the cabin), 1: low risk (no one but objects in the cabin), 2: medium risk (adult male in the cabin); 3: high risk (elderly / children / women / disabled people in the cabin)), there is no need to use special labeling software for labeling. It is only necessary to create four folders, named 0, 1, 2, and 3 respectively, and classify the collected images into corresponding folders according to the risk level. That is, pictures of no one and no objects in the cabin are placed in the "0" folder, pictures of no one but objects in the cabin are placed in the "1" folder, pictures of adult males in the cabin are placed in the "2" folder, and pictures of elderly / children / women / disabled people in the cabin are placed in the "3" folder.
[0082] 2. This example does not classify images, but text. Therefore, we need to call the multimodal model to describe the images in folders 0-3 and generate corresponding training samples. Assume that A0, B0, and C0 are in the "0" folder. Similarly, A1, B1, and C1 are in the "1" folder. A2, B2, and C2 are in the "2" folder. A3, B3, and C3 are in the "3" folder, as follows:
[0083]
[0084]
[0085] By calling the multimodal large model to describe all files, we get the corresponding description texts TA0, TB0, TC0, TA1, TB1, TC1, TA2, TB2, TC2, TA3, TB3, and TC3. Then, we combine these description texts and classification values into training sample files, as follows:
[0086] Classification label (risk level) Description text 0 TA0 0 TB0 0 TC0 1 TA1 1 TB1 1 TC1 2 TA2 2 TB2 2 TC2 3 TA3 3 TB3 3 TC3
[0087] For example:
[0088] The sample "0 There are no people or objects in the car, and the elevator is currently on the 3rd floor." means that the image description is "There are no people or objects in the car, and the elevator is currently on the 3rd floor." The image classification is "0", that is, the risk level is "0".
[0089] The sample "1 This is a photo of the interior of an elevator car. There is no one there, but there are some items, which may be boxes." means that the picture description is "This is a photo of the interior of an elevator car. There is no one there, but there are some items, which may be boxes." The picture classification is "1", that is, the risk level is "1".
[0090] At this point, this embodiment has converted the image classification problem into a text classification problem. The multimodal large model called by the image description can adopt a commercial model or an open source model. However, the multimodal large model used in the training process and the inference process should be kept as consistent as possible to ensure the accuracy of the inference results.
[0091] 3. After the sample data is generated, the sample confusion is divided into training set, validation set and test set in proportion. The training set is used to train the risk assessment model, the validation set is used to preliminarily evaluate the training effect of the risk assessment model, and the test set is used to confirm the final effect of the risk assessment model.
[0092] 4. The risk assessment model is implemented using a text classification model. The model structure is as follows: Figure 7 As shown in the figure; first, the Embedding layer uses 100, that is, the words are embedded (mapped) into 100-dimensional vectors. This is because the output of the pre-trained word vector used is 100-dimensional; the input of the BiLSTM layer is 100-dimensional, and the hidden layer is 128-dimensional. Since bidirectional LSTM is used, the actual output is (128*2)-dimensional; the input and output of the BatchNorm1d layer are both 256-dimensional; the input and output of the ReLU activation function layer are also 256-dimensional; the input of the Linear layer is 256-dimensional, and the output is 4-dimensional (that is, the probabilities of 4 categories); the input of the LogSoftmax activation function layer is 4-dimensional, and the output is 1-dimensional (that is, the category with the highest probability).
[0093] 5. Input data into the model in batches, calculate the loss using cross_entropy for each batch, and perform backward propagation. Use the validation set to verify the model accuracy every 100 batches. After training, use the test set to test the model accuracy. Figure 8 As shown in the figure, in the forward process, all layers will participate, but in the backward process, the BiLSTM layer, BatchNorm1d layer and Linear layer will be adjusted according to the Loss (as shown in the red box in the figure), and other layers will not be adjusted. The Embedding layer will not be adjusted because it uses pre-trained word vectors.
[0094] 6. After the model training is completed and the accuracy reaches the design index, the model is deployed online to provide a calling interface for other systems (such as communication services). The model reasoning process is the forward process, such as Figure 9 shown.
[0095] 7. After the system is deployed, images of the elevator cabin are saved during each interface call, allowing for continuous collection of images of the elevator cabin after people are trapped. These images can be used for continuous iterative optimization of the model. After a certain number of images have been collected, step 1 is repeated to continuously iterate and optimize the risk assessment model, improving its accuracy and generalization capabilities.
[0096] 8. The front-end system calls the image analysis service interface to obtain the risk level assessment result, determines the priority based on the risk level assessment result, and determines whether to automatically place an order based on the front-end system configuration. If the order is automatically placed, the order process is automatically called to send the fault handling form to the maintenance personnel. If the order is not automatically placed, the order is placed after confirmation by the customer service staff.
[0097] In summary, this solution uses the image description capabilities of a large multimodal model to achieve the conversion from images to natural language, thereby converting the image classification problem into a text classification problem, which is more intuitive and simple. At the same time, the use of open source or commercial large multimodal models not only saves the training cost of the large model, but also ensures that the description accuracy of the model itself is within our acceptable range, and can also be continuously iterated and optimized along with the model.
[0098] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.
Claims
1. A system for assessing the risk level of entrapment in elevator stalls based on a multimodal large model, characterized by: The system includes a DTU terminal, a communication server, an image analysis module and an automatic order module, wherein the DTU terminal is in communication connection with the communication server, the communication server is in communication connection with the image analysis module, and the image analysis module is in communication connection with the automatic order module; The DTU terminal is used to collect elevator operation status data and fault data in real time and send them to the communication server in real time. At the same time, the DTU terminal can perform image acquisition according to the instructions of the communication server; The communication server is used to receive the elevator operation status data and fault data uploaded by the DTU terminal, and control the DTU terminal to collect the car image when receiving the elevator stop fault data; The image analysis module is used to evaluate the risk level of the elevator car image and determine the processing priority; The automatic ordering module is used to automatically send the fault handling order to the maintenance personnel.
2. The risk level assessment method for elevator entrapment based on a multimodal large model is characterized by: The steps include: S1: receiving data uploaded by the elevator through the communication server, and capturing images when elevator stop fault data is received; S2: The communication server sends a command to collect car images, and the DTU controls the camera in the car to take pictures and transmit them back to the communication server; S3: After receiving the photo of the cabin, the communication server sends it to the image analysis module, and the image analysis module uses the multimodal large model to obtain a description of the image; S4: The image analysis module obtains the image description and sends it to the risk assessment model for risk assessment to obtain the risk level; S5: The image analysis module determines the processing priority according to the risk level, and the higher the risk, the higher the priority; S6: The automatic ordering module determines whether to automatically place an order based on the configuration.
3. The multimodal large model-based risk assessment method for elevator stalls and entrapment risks according to claim 2 is characterized by: When the communication server in S1 receives the elevator stop fault data, it first checks whether the DTU terminal supports the image acquisition function. If it does, it starts to acquire images. If it does not support image acquisition, the customer service staff confirms the elevator stop situation.
4. The multimodal large model-based risk assessment method for elevator stalls and entrapment risks according to claim 2 is characterized by: The risk levels of the risk assessment model in S4 include 0 to 3, where 0 indicates no risk, meaning there is no one and no items in the elevator car; 1 indicates low risk, meaning there is no one but items in the elevator car; 2 indicates medium risk, meaning there is an adult male in the elevator car; and 3 indicates high risk, meaning there are elderly people, children, women, or disabled people in the elevator car.
5. The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model according to claim 2 is characterized by: The multimodal large model in S3 is used to describe the collected images inside the car, thereby obtaining information inside the car described using natural language, and the natural language is input into the risk assessment model.
6. The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model according to claim 5 is characterized by: After receiving the natural language description text output by the multimodal large model, the risk assessment model performs word segmentation on the text, that is, dividing a sentence into multiple word sequences, and sending the word sequences into the risk assessment model in sequence.
7. The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model according to claim 2 is characterized in that: The model training process of the risk assessment model in S4 includes the following steps: S41: Label the pictures, label the collected car pictures, and mark the risk level according to the situation of the car pictures; S42: multimodal large model description, calling the multimodal large model to describe each car image and save the corresponding description text; S43: Integrate the description text and the label, and integrate the description text and risk level of the image to form a training sample.
8. The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model according to claim 2 is characterized in that: The model reasoning process of the risk assessment model in S4 includes the following steps: S411: Receiving a picture, the image analysis module receives a car picture sent by the communication server; S412: describing the multimodal large model, calling the multimodal large model to describe the car image and outputting corresponding description text; S413: Output the description text to the risk assessment model for risk level assessment.
9. The method for assessing the risk level of trapped people in an elevator stall based on a multimodal large model according to claim 2 is characterized by: If the automatic order module in S6 places an order automatically, the order process is automatically called to send the fault handling order to the maintenance personnel. If the order is not placed automatically, the order is placed after confirmation by the customer service personnel.
Citation Information
Patent Citations
Elevator operation and people trapping risk early warning method and system based on deep learning
CN118270617A
Visual rescue elevator and visual rescue method for elevator
CN105923484A
System and method for remote monitoring of trapping of elevator passengers
CN106586751A