Implementation method and system of elevator safety monitoring AI glasses based on multi-modal large model

By using multimodal large-scale model-based elevator safety monitoring AI glasses, real-time collaborative diagnosis of elevator structural anomalies and human behavioral risks has been achieved, solving the problem that traditional systems cannot establish causal relationships and improving the safety and reliability of elevator operation.

CN122166635APending Publication Date: 2026-06-09XJ SCHINDLER XUCHANG ELEVATOR
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XJ SCHINDLER XUCHANG ELEVATOR
Filing Date
2026-02-05
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional elevator safety monitoring systems cannot effectively build a unified scenario model and cannot establish a causal relationship between sudden passenger postural imbalance and elevator equipment malfunction, making it difficult to trace the root cause and provide early warnings in complex risk scenarios.

Method used

An AI glasses for elevator safety monitoring based on a multimodal large model is adopted. The AI ​​glasses simultaneously collect high-definition video streams, spatial audio streams, and wearer motion data streams with timestamp alignment. Multi-scale visual feature extraction, audio event detection, and posture calculation are performed. The system integrates and judges the movement status of personnel and the health status of the elevator structure to generate graded early warning signals.

Benefits of technology

It enables multimodal real-time collaborative diagnosis of elevator structural anomalies and human behavioral risks, improves the ability to detect and diagnose latent equipment faults in complex acoustic and optical environments, and enhances the safety and reliability of elevator operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122166635A_ABST
    Figure CN122166635A_ABST
Patent Text Reader

Abstract

This application provides a method and system for implementing AI glasses for elevator safety monitoring based on a multimodal large model. In the elevator operating environment, the AI ​​glasses synchronously collect multi-source sensor data streams with timestamp alignment. Based on the motion posture characteristics and abnormal behavior characteristics of the people inside the car, the motion state of the people inside the car is fused and judged to obtain the motion imbalance characteristics of the people inside the car. The health status of the elevator structure is collaboratively diagnosed through abnormal opening and closing states of the elevator doors, abnormal audio characteristics of the elevator at various operating stages, and the audio characteristics of the people inside the car calling for help, to obtain the fault characteristics of the elevator structure. Based on the motion imbalance characteristics of the people inside the car and the fault characteristics of the elevator structure, a risk fusion warning is performed on the elevator operating status, thereby obtaining a graded warning signal for elevator operation. Based on the above scheme, multimodal real-time collaborative diagnosis of elevator structural anomalies and personnel behavioral risks can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of elevator safety technology, and more specifically, to a method and system for implementing AI glasses for elevator safety monitoring based on a multimodal large model. Background Technology

[0002] Elevator safety is a comprehensive system that ensures the reliable operation of vertical transportation equipment and prevents personal injury and equipment damage. Through mechanical and electrical protection devices, real-time monitoring systems, and standardized operation and maintenance, it addresses major risks such as elevator falls, passenger entrapment, door system malfunctions, and overshooting / undershooting, ensuring passenger safety throughout the entire process from calling the elevator, riding the elevator, to leaving. Currently, technology is developing towards intelligent, predictive maintenance, and proactive safety warnings.

[0003] Traditional elevator safety is a comprehensive system that ensures the reliable operation of vertical transportation equipment and prevents personal injury and equipment damage. It addresses major risks such as falls, entrapment, door system failures, and overshooting / undershooting through mechanical and electrical protection devices, real-time monitoring systems, and standardized operation and maintenance. It ensures the safety of passengers throughout the entire process from calling the elevator, riding the elevator, to leaving the elevator. Currently, technology is developing towards intelligent, predictive maintenance, and proactive safety warnings. Elevator safety generally relies on isolated sensor arrays deployed inside and outside the elevator car, such as fixed cameras, microphones, and vibration sensors. These systems operate independently, and their data acquisition clocks are not synchronized and their processing protocols are heterogeneous. This results in video analysis modules being able to identify behavioral anomalies within the visual range, audio modules being able to capture specific noise events, and equipment vibration data being limited to mechanical condition analysis. The alarm information generated by each system lacks effective correlation and calibration in time and space, making it impossible to build a unified scenario model to establish a causal relationship between a passenger's sudden postural imbalance and simultaneous equipment failures such as abnormal opening and closing of the car doors and abnormal whistling of the traction machine. When faced with complex, chain-like risk scenarios, such as falls caused by hidden equipment malfunctions or increased equipment operational risks due to personnel disturbances, traditional systems struggle to perform root cause analysis and early warning systems. Therefore, achieving multimodal, real-time, collaborative diagnosis of elevator structural anomalies and personnel behavioral risks has become a significant challenge for the industry. Summary of the Invention

[0004] This application provides a method and system for implementing AI glasses for elevator safety monitoring based on a multimodal large model, which can realize multimodal real-time collaborative diagnosis of elevator structural anomalies and human behavioral risks.

[0005] Firstly, this application provides a method for implementing AI glasses for elevator safety monitoring based on a multimodal large model, including: In the elevator operating environment, AI glasses are used to synchronously collect multi-source sensor data streams with timestamp alignment. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. Multi-scale visual feature extraction is performed on the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car. Audio event detection is performed on the spatial audio stream to obtain the abnormal audio features of the elevator at each operating stage and the distress call audio features of the people in the car. Attitude calculation is performed on the motion data stream to obtain the motion attitude features of the people in the car. Based on the motion posture characteristics and the abnormal behavior characteristics, the motion state of the people in the car is fused and judged to obtain the motion imbalance characteristics of the people in the car. The health status of the elevator structure is diagnosed collaboratively by the abnormal opening and closing state, all abnormal audio characteristics and the emergency call audio characteristics to obtain the fault characteristics of the elevator structure. Based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, a risk fusion early warning is performed on the elevator operation status, thereby obtaining a graded early warning signal for elevator operation.

[0006] In some embodiments, multi-scale visual feature extraction is performed on the high-definition video stream to obtain abnormal opening and closing states of elevator doors and abnormal behavioral characteristics of people inside the car, specifically including: A dual-path visual feature extraction network for elevator environments is constructed. The first path uses a high-resolution backbone network to segment and track the elevator door area, while the second path uses a lightweight backbone network to understand the behavior of the car panorama. In the first path, based on the real-time position and speed information of the elevator door in the high-definition video stream, combined with the elevator operating floor signal, it is determined whether there is any door opening behavior such as misalignment with the floor, repeated opening and closing, or abnormal lingering, and the abnormal opening and closing state of the elevator door is obtained. In the second path, a time-series action localization algorithm is used to identify whether there are abnormal behavior patterns in the high-definition video stream, such as people squatting quickly, falling, pushing, remaining still for a long time, or being overcrowded, so as to obtain the abnormal behavior characteristics of the people in the car.

[0007] In some embodiments, performing audio event detection on the spatial audio stream to obtain abnormal audio characteristics of the elevator at various operating stages and the distress call audio characteristics of people in the car specifically includes: Based on the Mel frequency cepstral coefficients and time-domain envelope features, acoustic event detection is performed on the spatial audio stream to identify mechanical noise during elevator start-up, constant speed, deceleration and stopping stages, and the noise is compared with a standard operating acoustic signature database to obtain abnormal audio features of the elevator in each operating stage. By detecting and separating human voice components through voice activity, and combining keyword recognition and emotional speech analysis models, the audio features of people calling for help inside the car are extracted from the spatial audio stream.

[0008] In some embodiments, performing attitude calculation on the motion data stream to obtain the motion attitude characteristics of the people inside the car specifically includes: The motion data stream is fused and denoised to calculate the three-dimensional attitude angle and linear acceleration of the wearer's head. Based on the three-dimensional attitude angle and linear acceleration, the movement trend of the body center of gravity of the people in the car is determined, and the attitude features used to characterize the body balance state are extracted, thereby obtaining the motion attitude features of the people in the car.

[0009] In some embodiments, the motion state of the occupants in the elevator car is fused and judged based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the occupants in the elevator car, specifically including: The acceleration peak and sway intensity in the motion posture features are time-sequentially aligned with the falling and squatting actions in the abnormal behavior features. Feature fusion is performed on the temporally aligned motion posture features and abnormal behavior features to enhance the representation of the imbalance state of people in the car; The motion state of the people in the car at the current moment is determined based on the fused feature vector, and then the motion imbalance characteristics of the people in the car are output.

[0010] In some embodiments, the elevator structure's health status is collaboratively diagnosed using the abnormal opening and closing states, all abnormal audio features, and the distress call audio features to obtain the elevator structure's fault characteristics, specifically including: By establishing a temporal causal correlation between the abnormal opening and closing state of the elevator door and the abnormal audio features that appear in the same time period, the fault type and confidence level of the elevator door structure can be obtained. By associating the distress call audio features with abnormal audio features that appear in the same time period, the emergency safety risk type and level of the elevator car structure can be obtained. The fault characteristics of the elevator structure are determined by the fault type and confidence level of the elevator door structure and the emergency safety risk type and level of the elevator car structure.

[0011] In some embodiments, risk fusion early warning of elevator operation status is performed based on the motion imbalance characteristics of the occupants in the car and the fault characteristics of the elevator structure, thereby obtaining a graded early warning signal for elevator operation, specifically including: A risk matrix is ​​constructed to combine and map the motion imbalance characteristics of the people in the car with the fault characteristics of the elevator structure; Based on the mapping results, a preliminary warning level is generated using a preset risk decision tree. Then, risk weighting is applied to the superimposed situation of simultaneous personnel danger and equipment failure to obtain graded warning signals for elevator operation.

[0012] Secondly, this application provides an implementation system for elevator safety monitoring AI glasses based on a multimodal large model, including: The acquisition module is used to synchronously acquire multi-source sensor data streams with timestamp alignment in the elevator operating environment using AI glasses. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. The processing module is used to extract multi-scale visual features from the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car; to detect audio events from the spatial audio stream to obtain the abnormal audio features of the elevator at each stage of operation and the audio features of the people calling for help from the people in the car; and to perform attitude calculation on the motion data stream to obtain the motion attitude features of the people in the car. The processing module is also used to fuse and judge the motion state of the people in the car based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the people in the car, and to perform a collaborative diagnosis of the health status of the elevator structure through the abnormal opening and closing state, all abnormal audio characteristics and the emergency call audio characteristics to obtain the fault characteristics of the elevator structure. The execution module is used to perform risk fusion early warning of the elevator operation status based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, and then obtain a graded early warning signal for elevator operation.

[0013] Thirdly, this application provides a computer device, which includes a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computer device executes the above-described implementation method of the elevator safety monitoring AI glasses based on a multimodal large model.

[0014] Fourthly, this application provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the aforementioned method for implementing the elevator safety monitoring AI glasses based on a multimodal large model.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: This application provides a method and system for implementing elevator safety monitoring AI glasses based on a multimodal large model. In the elevator operating environment, the AI ​​glasses synchronously collect multi-source sensor data streams with timestamp alignment. These multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. Multi-scale visual feature extraction is performed on the high-definition video stream to obtain abnormal opening and closing states of the elevator doors and abnormal behavioral characteristics of people inside the car. Audio event detection is performed on the spatial audio stream to obtain abnormal audio features of the elevator at various operating stages and distress call audio features of people inside the car. Attitude calculation is performed on the motion data stream to obtain the motion posture features of people inside the car. The motion posture features and abnormal behavioral features are fused and judged to obtain the motion imbalance features of people inside the car. The health status of the elevator structure is collaboratively diagnosed using the abnormal opening and closing states, all abnormal audio features, and distress call audio features to obtain the fault features of the elevator structure. Based on the motion imbalance features of people inside the car and the fault features of the elevator structure, a risk fusion warning is performed on the elevator operating status to obtain a graded warning signal for elevator operation.

[0016] Therefore, this application utilizes the motion imbalance characteristics of the occupants in the elevator car and the fault characteristics of the elevator structure to perform risk fusion early warning of the elevator's operating status, thereby obtaining a graded early warning signal for elevator operation. Firstly, determining the motion posture characteristics yields a three-dimensional quantitative description of the wearer's posture within the elevator car, providing core ontological motion reference data for fusion judgment. This data is then processed through real-time data streams from the inertial measurement unit, eliminating environmental visual interference and directly characterizing the occupants' balance state, movement amplitude, frequency, and other intrinsic physical quantities. This allows for objective and continuous perception of human dynamics changes from a first-person perspective, laying a data foundation for accurately identifying motion imbalance risks and avoiding potential visual errors associated with solely relying on external video analysis. The system addresses issues such as corner occlusion and misjudgment. Then, by determining the fault characteristics of the elevator structure, a multi-dimensional integrated judgment of elevator mechanical and operational anomalies can be obtained, thereby achieving a comprehensive assessment of equipment health. It collaboratively utilizes visual recognition of door system anomalies, audio detection of operational abnormalities, and potential acoustic distress signals to construct a cross-modal fault diagnosis matrix. This enables the system to correlate and enhance dispersed and potentially weak anomaly indicators, thus more reliably identifying complex risks such as door lock malfunctions, guide rail noises, or entrapment incidents. This improves the early detection and diagnosis capability of latent equipment faults in complex acoustic and visual environments. In summary, based on the above scheme, multi-modal real-time collaborative diagnosis of elevator structural anomalies and personnel behavioral risks can be achieved. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an exemplary flowchart illustrating an implementation method of an AI glasses for elevator safety monitoring based on a multimodal large model, according to some embodiments of this application. Figure 2 This is a flowchart illustrating the process of determining fault characteristics according to some embodiments of this application; Figure 3 This is a schematic diagram of the implementation system of an elevator safety monitoring AI glasses based on a multimodal large model, according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a computer device that implements a method for using AI glasses for elevator safety monitoring based on a multimodal large model, according to some embodiments of this application. Detailed Implementation

[0019] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] refer to Figure 1 The figure is an exemplary flowchart illustrating an implementation method of an elevator safety monitoring AI glasses based on a multimodal large model, according to some embodiments of this application. The implementation method of the elevator safety monitoring AI glasses based on a multimodal large model mainly includes the following steps: In step 101, in the elevator operating environment, AI glasses are used to synchronously collect multi-source sensor data streams with timestamp alignment. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams.

[0021] It should be noted that, in this application, the AI ​​glasses integrate a forward-facing camera, a multi-microphone array, an inertial measurement unit, a near-eye display module, and a local computing module; the high-definition video stream is a first-person perspective visual data sequence used to acquire the elevator door status and personnel behavior; the spatial audio stream is an audio data sequence with directional information used to identify abnormal mechanical sounds and personnel cries for help; and the wearer motion data stream is an inertial measurement data sequence used to reflect the wearer's head posture and body balance.

[0022] In practice, the local computing module of the AI ​​glasses generates a unified time reference signal. The forward-facing camera captures first-person video covering the elevator door and the interior of the car at a fixed frame rate. The multi-microphone array synchronously captures audio data containing directional information. The inertial measurement unit continuously measures triaxial acceleration and triaxial angular velocity at a high sampling rate. In the data buffer of the AI ​​glasses, each frame of video, each audio data block, and each motion data sample is stamped with the same reference timestamp, thus obtaining a multi-source sensor data stream with precise time stamps.

[0023] In step 102, multi-scale visual feature extraction is performed on the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car. Audio event detection is performed on the spatial audio stream to obtain the abnormal audio features of the elevator at each operating stage and the distress call audio features of the people in the car. Attitude calculation is performed on the motion data stream to obtain the motion posture features of the people in the car.

[0024] In some embodiments, multi-scale visual feature extraction of the high-definition video stream to obtain abnormal opening and closing states of elevator doors and abnormal behavioral characteristics of people inside the car can be achieved through the following steps: A dual-path visual feature extraction network for elevator environments is constructed. The first path uses a high-resolution backbone network to segment and track the elevator door area, while the second path uses a lightweight backbone network to understand the behavior of the car panorama. In the first path, based on the real-time position and speed information of the elevator door in the high-definition video stream, combined with the elevator operating floor signal, it is determined whether there is any door opening behavior such as misalignment with the floor, repeated opening and closing, or abnormal lingering, and the abnormal opening and closing state of the elevator door is obtained. In the second path, a time-series action localization algorithm is used to identify whether there are abnormal behavior patterns in the high-definition video stream, such as people squatting quickly, falling, pushing, remaining still for a long time, or being overcrowded, so as to obtain the abnormal behavior characteristics of the people in the car.

[0025] It should be noted that, in this application, abnormal behavior features are feature representations used to quantitatively describe whether abnormal behavior patterns of personnel occur and their spatiotemporal context; the dual-path visual feature extraction network is a computational architecture used to perform refined region analysis and global scene understanding in parallel from a single video stream; the high-resolution backbone network is a deep neural network model used to preserve image details for pixel-level prediction; the lightweight backbone network is a deep neural network model used to quickly extract global semantic features under limited computing resources; and abnormal opening and closing states are categories of abnormal situations that characterize the operation of elevator doors in violation of preset safety or logical rules.

[0026] In specific implementation, firstly, a dual-path visual feature extraction network for the elevator environment is constructed. The first path uses a high-resolution backbone network to segment and track the elevator door area, while the second path uses a lightweight backbone network to understand the behavior of the elevator car panorama. This can be achieved as follows: After receiving the high-definition video stream, it is simultaneously input into two parallel neural network processing paths. In the first path, a specially designed or selected backbone network capable of maintaining high-resolution feature maps and performing upsampling operations is used to semantically segment the elevator door area in each frame, generating accurate door pixel masks. The displacement of the door is calculated using the mask changes between adjacent frames, achieving continuous segmentation and motion tracking of the door area. In the second path, a lightweight backbone network with fewer computational layers and smaller convolutional kernels is used to quickly process the complete video frames, extracting overall scene feature maps covering key targets such as people, handrails, and floor displays. The two network paths perform forward propagation calculations independently. Finally, the door area segmentation and tracking results output by the first path and the panoramic feature map output by the second path are used together as inputs for subsequent steps. Then, in the first path... In this process, based on the real-time position and speed information of the elevator door in the high-definition video stream, combined with the elevator floor operation signal, it is determined whether there is any door misalignment, repeated opening and closing, or abnormal lingering behavior. The abnormal opening and closing state of the elevator door can be obtained in the following way: The elevator door area tracking results output by the first path are processed, and the real-time horizontal or vertical movement speed of the elevator door is calculated according to the centroid change of the door pixel mask in continuous video frames. At the same time, the elevator floor operation signal obtained from the elevator control system or through the visual recognition floor display is received. A series of logical judgment rules are pre-set, namely: by comparing the visual alignment degree between the edge of the door and the car threshold after the door is fully opened, combined with the known floor plane position, it is determined whether "door misalignment" has occurred; by detecting whether the number of opening-closing-opening cycles of the door in a short period of time exceeds a threshold, it is determined whether "repeated opening and closing" has occurred; by calculating whether the duration of the door from the start of opening to full opening, or from full opening to the start of closing exceeds the normal time range, it is determined whether "abnormal lingering" has occurred; and the judgment conclusions that meet any one or more of the above rules are aggregated as the abnormal opening and closing state of the elevator door.Finally, in the second path, the temporal action localization algorithm is used to identify whether there are abnormal behavioral patterns such as rapid squatting, falling, pushing, prolonged stillness, or excessive crowding in the high-definition video stream. The abnormal behavior characteristics of the people in the car can be obtained in the following way: the temporal action localization algorithm is applied to the continuous multi-frame panoramic feature map sequence output by the second path. The algorithm first performs personnel target detection on each frame feature map, obtains the bounding box of each person and its basic posture key points. The algorithm analyzes the behavior of each tracked person along the time axis, and inputs the features such as the person's posture, movement trajectory, and relative distance between people in the continuous frames into a pre-trained classification model. The model can identify temporal features that conform to the predefined patterns such as "rapid squatting", "falling", and "pushing". For example, for the "falling" pattern, the model looks for a combination of features where the height of key points on the human torso drops sharply within a short period of time, accompanied by horizontal displacement; for "overcrowding," the model calculates whether the number of people per unit area exceeds a safety threshold, and the algorithm outputs a set of labels for each identified abnormal behavior pattern, the start and end times of the occurrence, and the identities of the people involved as the abnormal behavior characteristics of the people in the elevator car.

[0027] In some embodiments, performing audio event detection on the spatial audio stream to obtain abnormal audio characteristics of the elevator at various operating stages and the audio characteristics of people calling for help inside the car can be achieved by the following steps: Based on the Mel frequency cepstral coefficients and time-domain envelope features, acoustic event detection is performed on the spatial audio stream to identify mechanical noise during elevator start-up, constant speed, deceleration and stopping stages, and the noise is compared with a standard operating acoustic signature database to obtain abnormal audio features of the elevator in each operating stage. By detecting and separating human voice components through voice activity, and combining keyword recognition and emotional speech analysis models, the audio features of people calling for help inside the car are extracted from the spatial audio stream.

[0028] It should be noted that, in this application, abnormal audio features are quantitative descriptions of the deviation of elevator mechanical noise from the normal pattern; distress call audio features are composite features that characterize whether there is distress semantics and high urgency in the speech.

[0029] In specific implementation, firstly, acoustic event detection is performed on the spatial audio stream based on Mel-frequency cepstral coefficients and temporal envelope features to identify mechanical noise during elevator startup, constant speed, deceleration, and stopping phases. This noise is then compared with a standard operating acoustic signature database to obtain abnormal audio features of the elevator at each operating phase. This can be achieved through the following methods: preprocessing the spatial audio stream, including pre-emphasis, framing, and windowing; calculating the Mel-frequency cepstral coefficients for each frame of audio. This calculation process includes performing a Fast Fourier Transform (FFT) on the signal to obtain the spectrum, smoothing and reducing the dimensionality of the spectrum using a Mel-scale filter bank simulating the critical bandwidth of the human ear, and then performing a Discrete Cosine Transform (DCT) on the logarithm of the filter bank's output to obtain the cepstral coefficients representing the shape of the spectral envelope; simultaneously, short-time energy, zero-crossing rate, and other temporal envelope features are extracted from the original audio frames to supplement the transient and energy information of the audio. The extracted Mel-frequency cepstral coefficients and temporal envelope features are combined into a joint feature vector, which is then input into a pre-trained acoustic event detection classification model. The model can map the joint feature vector to categories such as "start", "constant speed", "deceleration" and "stop" of elevator operation, and identify the mechanical noise components. It compares the identified current mechanical noise features with the normal voiceprint features of the corresponding operation stage in the standard operation voiceprint library and calculates their distance or difference.If the difference exceeds a preset abnormal threshold, it is judged as abnormal. The information including the abnormal judgment result, the operating stage, and the specific difference value is combined as the abnormal audio features of the elevator in each operating stage. Then, human voice components are separated through voice activity detection, and combined with keyword recognition and emotional speech analysis models, the audio features of people calling for help in the car can be extracted from the spatial audio stream in the following way: A voice activity detection algorithm is applied to the entire spatial audio stream. This algorithm is usually based on the statistical characteristics of short-time energy and spectral features (e.g., Mel frequency cepstral coefficients), setting a dynamic threshold to determine whether each frame of audio contains valid human voice components, thereby dividing the audio stream into speech segments and non-speech segments. The speech segments are then extracted, and noise reduction and enhancement processing is performed on the extracted speech segments to obtain relatively pure human voice components. The processed speech data is then sent to two analysis processes in parallel: In the first process, a... A pre-trained keyword recognition model scans the entire speech segment. This model, based on a deep neural network, can identify predefined distress call-related keywords such as "help me," "help me," and "stop," along with the timing of their occurrence. In the second process, an emotional speech analysis model analyzes the speech segment. This model determines whether the speaker is in a high-urgency emotional state, such as "fear," "tension," or "urgency," by analyzing acoustic parameters such as the fundamental frequency, formants, speech rate, and energy changes. The keyword list and its confidence level output by the keyword recognition model are combined with the emotional state category and intensity output by the emotional speech analysis model. For example, when both a high-confidence distress call keyword and a high-intensity urgency modality are detected simultaneously, a strong distress call feature is generated. The fused judgment result, including the presence of a distress call, confidence level, and emotional intensity, is used as the audio feature of the distress call from the people inside the elevator car.

[0030] In some embodiments, the attitude calculation of the motion data stream to obtain the motion attitude characteristics of the people inside the car can be achieved by the following steps: The motion data stream is fused and denoised to calculate the three-dimensional attitude angle and linear acceleration of the wearer's head. Based on the three-dimensional attitude angle and linear acceleration, the movement trend of the body center of gravity of the people in the car is determined, and the attitude features used to characterize the body balance state are extracted, thereby obtaining the motion attitude features of the people in the car.

[0031] It should be noted that, in this application, the three-dimensional attitude angle is an Euler angle representation used to describe the rotation angle of an object in three-dimensional space around the three axes of its own coordinate system (usually pitch angle, roll angle, and yaw angle); and the linear acceleration is a vector used to describe the rate of change of the translational velocity of an object in three-dimensional space along the three axes of its own coordinate system.

[0032] In practical implementation, firstly, the motion data stream is fused and denoised to calculate the wearer's three-dimensional head posture angle and linear acceleration. This can be achieved by receiving the raw data streams from the AI ​​glasses' inertial measurement unit, including the three-axis accelerometer, three-axis gyroscope, and three-axis magnetometer. A fusion calculation method based on gradient descent and complementary filtering principles is used. The core of this method lies in establishing a mathematical model of head posture represented by quaternions and using the integral of the angular velocity measured by the gyroscope to predict posture changes. Simultaneously, the specific force vector measured by the accelerometer (primarily reflecting the direction of gravity under static or quasi-static conditions) is used to correct the accumulated errors in attitude prediction on the pitch and roll axes, and the direction of the geomagnetic field measured by the magnetometer is used to correct the accumulated errors on the yaw axis (heading angle). The filter adaptively adjusts the weighting coefficients between gyroscope data and accelerometer / magnetometer correction data to achieve optimal fusion estimation under dynamic conditions, effectively suppressing gyroscope drift noise and accelerometer vibration noise. Through this fusion algorithm, the precise head orientation, represented in quaternion form, can be calculated in real time and then converted into easily understandable pitch. The angles are roll angle, yaw angle, and roll angle, i.e., three-dimensional attitude angles. At the same time, the gravity component calculated from the current attitude is subtracted from the original three-axis accelerometer data to obtain the linear acceleration reflecting the pure translational motion of the head. The calculated three-dimensional attitude angle vector and the linear acceleration vector are used as the output of this step. Then, based on the three-dimensional attitude angle and linear acceleration, the trend of the body center of gravity movement of the people in the car is determined, and the attitude features used to characterize the body balance state are extracted. The motion attitude features of the people in the car can be obtained by the following method: establishing an approximate mapping model from head movement to body center of gravity movement.This model, based on the biomechanical characteristics of human movement, posits a strong correlation between head movement and the movement of the torso and center of gravity during normal standing or movement. By analyzing the time-series data of the calculated three-dimensional posture angles, it calculates their rate of change (angular velocity) and quadratic rate of change (angular acceleration), particularly the changes in pitch and roll angles, which are directly related to the body's forward / backward and left / right tilt. Simultaneously, it analyzes the magnitude and direction of the horizontal components of linear acceleration (forward / backward and left / right), performing correlation analysis between head angular velocity, angular acceleration, and horizontal linear acceleration within a time window. For example, a large forward linear acceleration accompanied by... A continuously increasing forward pitch velocity of the head may indicate that the body's center of gravity is accelerating forward and has a tendency to lean forward. Based on these analyses, a set of predefined quantitative indicators are calculated, such as: the standard deviation of the posture angle (reflecting the degree of swaying), the peak value of the angular velocity (reflecting the intensity of the posture change), the energy of the linear acceleration in a specified frequency band (reflecting the characteristics of a specified type of vibration or gait), and the matching score between head movement and typical imbalance patterns (e.g., lateral swaying, forward leaning, backward leaning). The calculated quantitative indicators are then combined into a multi-dimensional feature vector as a comprehensive characterization of the motion posture and balance state of the person in the car.

[0033] In step 103, the motion state of the people in the car is fused and judged based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the people in the car. The health status of the elevator structure is diagnosed collaboratively by the abnormal opening and closing state, all abnormal audio characteristics and the distress call audio characteristics to obtain the fault characteristics of the elevator structure.

[0034] In some embodiments, the motion state of the people inside the car is fused and judged based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the people inside the car, which can be achieved by the following steps: The acceleration peak and sway intensity in the motion posture features are time-sequentially aligned with the falling and squatting actions in the abnormal behavior features. Feature fusion is performed on the temporally aligned motion posture features and abnormal behavior features to enhance the representation of the imbalance state of people in the car; The motion state of the people in the car at the current moment is determined based on the fused feature vector, and then the motion imbalance characteristics of the people in the car are output.

[0035] It should be noted that, in this application, motion imbalance features are features used to indicate that the people in the car are in an unbalanced state; peak acceleration is used to describe the maximum instantaneous value of the linear acceleration vector amplitude within a specified time window; sway intensity is a statistical measure used to quantify the degree of drastic change in head posture angle within a specified time window.

[0036] In specific implementation, firstly, the timing alignment of the acceleration peak and sway intensity in the motion posture features with the fall and squatting actions in the abnormal behavior features can be achieved as follows: Acquire two sets of time-series data: motion posture features and abnormal behavior features. Since the data sources (inertial measurement unit and visual analysis module) for the two features may have different initial timestamps and sampling frequencies, time synchronization is required. Using a unified world time as the reference, and utilizing the timestamps injected during the acquisition phase, the sequence of acceleration peak and sway intensity in the motion posture features is aligned with the sequences of "fall start," "squatting start," and "action" identified in the abnormal behavior features through linear interpolation or nearest neighbor matching. The time points and duration information of events such as "continuous" are mapped to the same high-precision timeline. Specifically, for each abnormal behavior event identified by the visual analysis module, the corresponding precise start and end times are found on the timeline, and all overlapping or adjacent motion posture feature data points within that time period are extracted. This ensures a strict correspondence between the two in the time dimension. For example, when the visual module determines that a "fall" event occurs between time T0 and T1, it extracts all acceleration peaks and sway intensity data from a short time window before T0 to the end of a short time window after T1 from the motion posture feature sequence. These motion data are then associated with the "fall" event. A timeline can be generated in this way. A data structure containing the correlation between time-aligned motion features and visual behavioral events is provided. Then, feature fusion is performed on the time-aligned motion posture features and abnormal behavior features to enhance the representation of the imbalance state of people inside the car. This can be achieved in the following way: The time-aligned motion posture features are received; to achieve effective feature fusion, a multimodal fusion network is constructed or a feature splicing and transformation strategy is adopted; feature engineering processing is performed on the time-aligned data: for motion posture features (e.g., acceleration peak sequence, sway intensity sequence), their statistical characteristics (e.g., mean, variance, maximum value, minimum value, energy of a specified frequency band) within the corresponding abnormal behavior event window may be calculated; for abnormal behavior... Features, in addition to event category labels (falling, squatting, etc.), may also include quantitative parameters obtained from visual analysis, such as the depth of squatting and the complexity of the trajectory of key points of the human body during the fall. These features from different modalities, which have been preliminarily summarized, are concatenated to form a wide-dimensional original fused feature vector. In order to enhance the representation ability and suppress redundancy, this original fused feature vector can be input into a lightweight feature transformation module, such as a neural network with one or two fully connected layers. This transformation module performs nonlinear transformation and dimensionality reduction on the original concatenated features through the learned weight matrix, and extracts high-level abstract features that can more essentially and compactly reflect the commonalities and complementarities of multimodal information.This process outputs a fused and enhanced feature vector of appropriate dimensionality. This feature vector integrates evidence from both motion and vision, aiming to more robustly represent the potential imbalance state of the occupants in the elevator car. Finally, the motion state of the occupants in the elevator car at the current moment is determined based on the fused feature vector, and the motion imbalance feature of the occupants in the elevator car is output. This can be achieved in the following way: the fused feature vector is used as input and fed into a pre-trained motion state classifier. This classifier can adopt model structures such as logistic regression, support vector machine, or shallow neural network. The output layer of the classifier corresponds to multiple preset motion state categories, such as "stable", "slight imbalance", "about to fall", and "already fallen". During the inference process, the classifier calculates the probability score or decision function value of the input feature vector belonging to each category according to the learned decision boundary or probability distribution. The category with the highest probability score is selected as the final judgment of the motion state of the occupants in the elevator car at the current moment, and this category is determined as the motion imbalance label. At the same time, the highest probability score itself (or the score after normalization) is used as the confidence level of this judgment. For example, the classifier might output a probability of 0.92 for "has fallen," while the probabilities for other categories are very low. In this case, the imbalance label would be "has fallen," with a confidence level of 0.92. The judgment result, containing the specific label and the high / low confidence level, is then encapsulated as the imbalance feature of the person inside the elevator car.

[0037] In some embodiments, the health status of the elevator structure is collaboratively diagnosed using the abnormal opening and closing states, all abnormal audio features, and the distress call audio features to obtain the fault characteristics of the elevator structure, for reference. Figure 2 The diagram is a flowchart illustrating the process of determining fault characteristics in some embodiments of this application. In this embodiment, the determination of fault characteristics can be achieved using the following steps: In step 1031, the abnormal opening and closing state of the elevator door is correlated with the abnormal audio features that appear in the same time period to obtain the fault type and confidence level of the elevator door structure. In step 1032, the distress call audio features are correlated with abnormal audio features that appear in the same time period to obtain the emergency safety risk type and level of the elevator car structure. In step 1033, the fault characteristics of the elevator structure are determined by the fault type and confidence level of the elevator door structure and the emergency safety risk type and level of the elevator car structure.

[0038] It should be noted that, in this application, the fault characteristics of the elevator structure are comprehensive diagnostic results used to structurally describe the current equipment faults and safety risks of the elevator system; the fault type and confidence level are combined information used to describe the possible failure modes of the elevator door structure and the reliability of their judgment results; and the emergency safety risk type and level are assessment results used to describe the specific hazard categories and their severity that the car structure faces and that require immediate attention.

[0039] In practical implementation, firstly, the abnormal opening and closing states of the elevator doors are correlated temporally and causally with abnormal audio features occurring in the same time period to obtain the fault type and confidence level of the elevator door structure. This can be achieved by establishing a pre-defined elevator door fault knowledge base, which defines the possible temporal and causal correspondences between different "abnormal opening and closing states" (e.g., door jamming, repeated opening and closing, misalignment with floors) and specified "abnormal audio features" (e.g., sharp friction sound, intermittent impact sound, continuous vibration sound). For example, the knowledge base may contain a rule: if a "door jamming" event is accompanied by "sharp friction sound in the guide rail area" within the occurrence time window... "Sharp friction sound" and "abnormal opening / closing state" are highly correlated, both pointing to the fault type of "foreign object or deformation in the door guide rail". In specific operation, for each detected "abnormal opening / closing state" event, the system extends a preset time window before and after its occurrence, retrieving all identified "abnormal audio features" within this window. The retrieved audio features are then matched with possible audio patterns corresponding to that abnormal opening / closing state in the knowledge base. The matching process considers not only whether the feature categories match, but also their temporal overlap, sequence (e.g., whether the abnormal noise occurs at the moment the door movement is obstructed), and intensity correlation. For each... For each matched knowledge base rule, an initial confidence score is calculated based on the degree of matching (e.g., time overlap, feature similarity). All matching results are summarized, and the fault type pointed to by the rule with the highest initial confidence score is selected as the most likely diagnosis, with that confidence score serving as its confidence level. The output is the fault type and confidence level of the elevator door structure. Then, the emergency call audio features are correlated with abnormal audio features occurring in the same time period to obtain the emergency safety risk type and level of the elevator car structure. This can be achieved by establishing a preset risk scenario rule base, which defines... When "distress call audio features" (e.g., detected distress keywords, high urgency) and various "abnormal audio features" (e.g., violent impact sounds, metal cracking sounds, abnormal emergency stop sounds) appear in combination under the specified scenario of elevator operation, the potential emergency safety risks implied are considered. For example, a rule might stipulate that if "distress call audio features" (indicating panic) and "violent impact sounds at the bottom of the car" occur almost simultaneously, it may indicate a risk of "unexpected car movement or sudden stop causing passenger injury"; if "distress call audio features" are accompanied by "continuous abnormal vibration and scraping sounds," it may indicate "car jamming or entrapment." Specifically, for each detected "distress call audio feature" instance, the system defines a correlation time window near its timestamp and searches for all "abnormal audio features" appearing within that window. The combination of these features is compared with entries in the risk scenario rule base; the comparison not only checks whether feature categories occur simultaneously but also assesses the correlation between the emotional intensity of the distress call and the physical intensity of the abnormal sound.For successfully matched rules, a risk level score is calculated from a pre-defined mapping table based on the confidence level of the distress call features, the intensity and typicality of the abnormal audio features, and the temporal proximity between events. This score may be quantified as "low," "medium," "high," or a specific numerical level. The risk type corresponding to the successfully matched rule (e.g., "passenger injured due to sudden car stop") and its calculated risk level are then used together as the emergency safety risk type and level output for the elevator car structure. Finally, this can be achieved by taking the fault type and confidence level of the elevator door structure and the risk assessment results of the car structure as input, and executing a comprehensive decision-making process to generate a unified fault feature representation. This process first assesses the emergency safety risk level of the car structure; if the risk level is "high," it indicates a direct risk. For threats to personal safety or major equipment hazards, priority will be given to highlighting the risk of the car structure, making it a primary component of the current elevator structure's fault characteristics, while retaining the diagnostic results of the door structure as secondary or related information. If the risk level of the car structure is "medium" or "low," the diagnostic confidence of the door structure will be further examined. If the diagnostic confidence of the door structure fault is higher than the preset high confidence threshold, the door structure fault will be taken as the primary feature. Conversely, if neither is prominent, information fusion will be performed to generate a composite description. The primary fault / risk information, secondary information, their respective types, levels / confidence, and possible correlations between them (e.g., door faults may lead to abnormal car operation and thus trigger risks) will be integrated into a structured data object or vector as the elevator structure's fault characteristics.

[0040] In step 104, risk fusion early warning is performed on the elevator operation status based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, thereby obtaining a graded early warning signal for elevator operation.

[0041] In some embodiments, the risk fusion warning of elevator operation status based on the motion imbalance characteristics of the occupants in the car and the fault characteristics of the elevator structure, and thus obtaining a graded warning signal for elevator operation, can be achieved through the following steps: A risk matrix is ​​constructed to combine and map the motion imbalance characteristics of the people in the car with the fault characteristics of the elevator structure; Based on the mapping results, a preliminary warning level is generated using a preset risk decision tree. Then, risk weighting is applied to the superimposed situation of simultaneous personnel danger and equipment failure to obtain graded warning signals for elevator operation.

[0042] In practical implementation, firstly, a risk matrix is ​​constructed. The combination and mapping of the motion imbalance characteristics of the occupants in the elevator car with the fault characteristics of the elevator structure can be achieved as follows: A two-dimensional risk matrix is ​​designed and constructed. One dimension (e.g., rows) represents the severity of the motion imbalance of the occupants in the car, derived from the motion imbalance characteristics, typically quantified as discrete levels such as "stable," "slightly unbalanced," "about to fall," and "already fallen." The other dimension (e.g., columns) represents the severity of the elevator structure fault characteristics, integrating fault type, confidence level, and emergency safety risk. The risk level is also quantified into discrete levels, such as "normal," "minor fault," "moderate fault," and "serious fault / high risk." Each cell in the risk matrix is ​​predefined with a comprehensive risk level (e.g., "Level I (low)," "Level II (medium)," "Level III (high)," and "Level IV (emergency)"). In practical applications, upon receiving real-time motion imbalance features (including imbalance labels) and elevator structural fault features (including comprehensive severity judgment), the corresponding levels in their respective dimensions are determined based on these two input values. Then, using row and column levels as indices, the corresponding cell in the risk matrix is ​​searched, and the predefined comprehensive risk level of that cell is used as the output of this combination mapping, thus completing the quantification conversion from specific features to comprehensive risk levels. Based on the mapping results, a pre-set risk decision tree is applied to generate preliminary warning levels. Then, risk weighting is applied to cases where both personnel danger and equipment failure exist simultaneously, resulting in a graded warning signal for elevator operation. This can be achieved by using the comprehensive risk level as the root node input of the risk decision tree, which consists of a series of pre-set logical judgment rules. For example, the rule setting is: if the comprehensive risk level is equal to... If the level is "Level I (Low)", the initial warning level is "Caution"; if it is "Level II (Medium)", it is "Warning"; if it is "Level III (High)", it is "Alarm"; if it is "Level IV (Emergency)", it is "Emergency Alarm". This constitutes the initial warning level. Perform superimposed situation detection: check whether the current input feature simultaneously meets the "Personnel Hazard" condition (e.g., the movement imbalance label is "about to fall" or "has fallen") and the "Equipment Failure" condition (e.g., the severity of the elevator structural failure feature is "Moderate Failure" or higher). If both are met, the risk weighted escalation rule is triggered. The rule may stipulate that, in this superimposed situation, the initial warning level will be upgraded by at least one level (e.g., from "warning" to "alarm", or from "alarm" to "emergency alarm"). The upgrade rule can be further refined, for example, by upgrading to different levels based on the specific type of equipment failure (e.g., whether it involves operational safety) or the immediacy of personnel danger; thereby encapsulating the warning level after weighted upgrade adjustment (or maintaining the original initial level if no upgrade is triggered) into a tiered warning signal output.

[0043] Furthermore, in another aspect of this application, in some embodiments, this application provides an implementation system for elevator safety monitoring AI glasses based on a multimodal large model, referencing... Figure 3 The figure is a schematic diagram of the structure of an elevator safety monitoring AI glasses system based on a multimodal large model, according to some embodiments of this application. The system includes a data acquisition module 201, a processing module 202, and an execution module 203, which are described below: The acquisition module 201 in this application is mainly used to synchronously acquire multi-source sensor data streams with timestamp alignment using AI glasses in the elevator operating environment. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. Processing module 202, in this application, is used to extract multi-scale visual features from the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car; to detect audio events from the spatial audio stream to obtain the abnormal audio features of the elevator at each operating stage and the distress call audio features of the people in the car; and to perform attitude calculation on the motion data stream to obtain the motion attitude features of the people in the car. It should be noted that the processing module 202 is also used to perform a fusion judgment on the motion state of the people in the car based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the people in the car, and to perform a collaborative diagnosis on the health status of the elevator structure through the abnormal opening and closing state, all abnormal audio characteristics and the emergency call audio characteristics to obtain the fault characteristics of the elevator structure. The execution module 203 in this application is mainly used to perform risk fusion early warning of the elevator operation status based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, and then obtain the graded early warning signal of the elevator operation.

[0044] The foregoing detailed examples of the implementation method and system of the elevator safety monitoring AI glasses based on a multimodal large model provided in this application. It is understood that the corresponding device, in order to achieve the above functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specified application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specified application, but such implementation should not be considered beyond the scope of this application.

[0045] In some embodiments, this application also provides a computer device, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, so that the computer device executes the above-described implementation method of the elevator safety monitoring AI glasses based on a multimodal large model.

[0046] In some embodiments, reference Figure 4 The dashed lines in the figure indicate that the unit or module is optional. This figure is a schematic diagram of the structure of a computer device for implementing an elevator safety monitoring AI glasses based on a multimodal large model, according to an embodiment of this application. The implementation method of the elevator safety monitoring AI glasses based on a multimodal large model described in the above embodiments can be achieved through… Figure 4 The computer device shown is used to implement this, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device may be a terminal device, a server or a chip.

[0047] Processor 301 can be a general-purpose processor or a special-purpose processor. For example, processor 301 can be a central processing unit (CPU), which can be used to control computer devices, execute software programs, and process data from software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.

[0048] For example, the computer device may be a chip, and the communication unit 305 may be the input and / or output circuit of the chip, or the communication unit 305 may be the communication interface of the chip, which may be a component of a terminal device, network device or other device.

[0049] For example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.

[0050] The computer device may include one or more memories 302 storing a program 304. The program 304 can be executed by a processor 301 to generate instructions 303, causing the processor 301 to execute the method described in the above method embodiments according to the instructions 303. Optionally, the memory 302 may also store data (such as a target audit model). Optionally, the processor 301 may also read data stored in the memory 302, which may be stored at the same storage address as the program 304, or the data may be stored at a different storage address than the program 304.

[0051] The processor 301 and memory 302 can be configured separately or integrated together, for example, integrated on the system on chip (SOC) of the terminal device.

[0052] It should be understood that each step of the above method embodiment can be completed by hardware logic circuits or software instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.

[0053] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0054] For example, in some embodiments, this application also provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the above-described method for implementing the elevator safety monitoring AI glasses based on a multimodal large model.

[0055] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0056] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for implementing AI glasses for elevator safety monitoring based on a multimodal large model, characterized in that, Includes the following steps: In the elevator operating environment, AI glasses are used to synchronously collect multi-source sensor data streams with timestamp alignment. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. Multi-scale visual feature extraction is performed on the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car. Audio event detection is performed on the spatial audio stream to obtain the abnormal audio features of the elevator at each operating stage and the distress call audio features of the people in the car. Attitude calculation is performed on the motion data stream to obtain the motion attitude features of the people in the car. Based on the motion posture characteristics and the abnormal behavior characteristics, the motion state of the people in the car is fused and judged to obtain the motion imbalance characteristics of the people in the car. The health status of the elevator structure is diagnosed collaboratively by the abnormal opening and closing state, all abnormal audio characteristics and the emergency call audio characteristics to obtain the fault characteristics of the elevator structure. Based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, a risk fusion early warning is performed on the elevator operation status, thereby obtaining a graded early warning signal for elevator operation.

2. The method as described in claim 1, characterized in that, Multi-scale visual feature extraction is performed on the high-definition video stream to obtain the abnormal opening and closing status of the elevator doors and the abnormal behavior characteristics of the people inside the car, specifically including: A dual-path visual feature extraction network for elevator environments is constructed. The first path uses a high-resolution backbone network to segment and track the elevator door area, while the second path uses a lightweight backbone network to understand the behavior of the car panorama. In the first path, based on the real-time position and speed information of the elevator door in the high-definition video stream, combined with the elevator operating floor signal, it is determined whether there is any door opening behavior such as misalignment with the floor, repeated opening and closing, or abnormal lingering, and the abnormal opening and closing state of the elevator door is obtained. In the second path, a time-series action localization algorithm is used to identify whether there are abnormal behavior patterns in the high-definition video stream, such as people squatting quickly, falling, pushing, remaining still for a long time, or being overcrowded, so as to obtain the abnormal behavior characteristics of the people in the car.

3. The method as described in claim 1, characterized in that, Audio event detection is performed on the spatial audio stream to obtain abnormal audio characteristics of the elevator at various operating stages and the audio characteristics of people calling for help inside the car. Specifically, this includes: Based on the Mel frequency cepstral coefficients and time-domain envelope features, acoustic event detection is performed on the spatial audio stream to identify mechanical noise during elevator start-up, constant speed, deceleration and stopping stages, and the noise is compared with a standard operating acoustic signature database to obtain abnormal audio features of the elevator in each operating stage. By detecting and separating human voice components through voice activity, and combining keyword recognition and emotional speech analysis models, the audio features of people calling for help inside the car are extracted from the spatial audio stream.

4. The method as described in claim 1, characterized in that, The attitude calculation of the motion data stream yields the motion attitude characteristics of the people inside the car, specifically including: The motion data stream is fused and denoised to calculate the three-dimensional attitude angle and linear acceleration of the wearer's head. Based on the three-dimensional attitude angle and linear acceleration, the movement trend of the body center of gravity of the people in the car is determined, and the attitude features used to characterize the body balance state are extracted, thereby obtaining the motion attitude features of the people in the car.

5. The method as described in claim 1, characterized in that, Based on the motion posture characteristics and the abnormal behavior characteristics, the motion state of the people in the car is fused and judged to obtain the motion imbalance characteristics of the people in the car, which specifically include: The acceleration peak and sway intensity in the motion posture features are time-sequentially aligned with the falling and squatting actions in the abnormal behavior features. Feature fusion is performed on the temporally aligned motion posture features and abnormal behavior features to enhance the representation of the imbalance state of people in the car; The motion state of the people in the car at the current moment is determined based on the fused feature vector, and then the motion imbalance characteristics of the people in the car are output.

6. The method as described in claim 1, characterized in that, By collaboratively diagnosing the health status of the elevator structure through the abnormal opening and closing states, all abnormal audio features, and the distress call audio features, the specific fault characteristics of the elevator structure are obtained, including: By establishing a temporal causal correlation between the abnormal opening and closing state of the elevator door and the abnormal audio features that appear in the same time period, the fault type and confidence level of the elevator door structure can be obtained. By associating the distress call audio features with abnormal audio features that appear in the same time period, the emergency safety risk type and level of the elevator car structure can be obtained. The fault characteristics of the elevator structure are determined by the fault type and confidence level of the elevator door structure and the emergency safety risk type and level of the elevator car structure.

7. The method as described in claim 1, characterized in that, Based on the motion imbalance characteristics of the occupants inside the elevator car and the fault characteristics of the elevator structure, a risk fusion early warning system is implemented for the elevator operating status, resulting in graded early warning signals for elevator operation. Specifically, these signals include: A risk matrix is ​​constructed to combine and map the motion imbalance characteristics of the people in the car with the fault characteristics of the elevator structure; Based on the mapping results, a preliminary warning level is generated using a preset risk decision tree. Then, risk weighting is applied to the superimposed situation of simultaneous personnel danger and equipment failure to obtain graded warning signals for elevator operation.

8. A system for implementing AI glasses for elevator safety monitoring based on a multimodal large model, characterized in that, include: The acquisition module is used to synchronously acquire multi-source sensor data streams with timestamp alignment in the elevator operating environment using AI glasses. The multi-source sensor data streams include high-definition video streams, spatial audio streams, and wearer motion data streams. The processing module is used to extract multi-scale visual features from the high-definition video stream to obtain the abnormal opening and closing state of the elevator door and the abnormal behavior features of the people in the car; to detect audio events from the spatial audio stream to obtain the abnormal audio features of the elevator at each stage of operation and the audio features of the people calling for help from the people in the car; and to perform attitude calculation on the motion data stream to obtain the motion attitude features of the people in the car. The processing module is also used to fuse and judge the motion state of the people in the car based on the motion posture characteristics and the abnormal behavior characteristics to obtain the motion imbalance characteristics of the people in the car, and to perform a collaborative diagnosis of the health status of the elevator structure through the abnormal opening and closing state, all abnormal audio characteristics and the emergency call audio characteristics to obtain the fault characteristics of the elevator structure. The execution module is used to perform risk fusion early warning of the elevator operation status based on the motion imbalance characteristics of the people in the car and the fault characteristics of the elevator structure, and then obtain a graded early warning signal for elevator operation.

9. A computer device, characterized in that, The computer device includes a memory and a processor. The memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that the computer device executes the implementation method of the elevator safety monitoring AI glasses based on a multimodal large model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions or code that, when executed on a computer, cause the computer to implement the method for implementing the elevator safety monitoring AI glasses based on a multimodal large model as described in any one of claims 1 to 7.