Method and device for monitoring state of personnel in elevator, terminal and medium

Through facial and full-body posture recognition, the preset model is used to calculate the status scores of personnel in the elevator, which solves the problem of inaccurate judgment of personnel in the elevator in the existing technology, and achieves rapid, accurate judgment and timely processing of personnel in the elevator.

CN120279474APending Publication Date: 2025-07-08SICHUAN SPECIAL EQUIP INSPECTION & RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149785.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing personnel monitoring methods in the elevator cannot quickly and accurately monitor and judge the status of the personnel in the elevator through the pictures or sensors captured by the camera, especially in a dense environment, and it is difficult to make accurate status judgments.

Method used

Through facial recognition and full-body posture recognition, the recognition scores of facial images and full-body images are obtained respectively by using the preset first model and second model, and the status scores of personnel in the elevator are calculated based on these scores to achieve accurate judgment of the status of personnel in the elevator.

Benefits of technology

It can quickly and accurately judge the status of personnel in the elevator, so as to deal with sudden abnormalities in a timely manner, improving the safety of the elevator and passenger comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279474A_ABST
    Figure CN120279474A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for monitoring the state of people in an elevator, a terminal and a medium. The method comprises the steps that face images and whole-body images of people in the elevator are obtained; obtaining a face recognition score according to the face image through a preset first model, and determining a first weight based on the face recognition score; obtaining a posture recognition score according to the whole body image through a preset second model, and determining a second weight based on the posture recognition score; and on the basis of the face recognition score, the posture recognition score, the first weight and the second weight, a state score corresponding to the person in the elevator is obtained through calculation, and on the basis of the state score, a target state monitoring result corresponding to the person in the elevator is determined. The invention aims to accurately judge the state of people in the elevator by integrating different angles through face recognition and whole body posture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of personnel monitoring, and particularly to a method, device, terminal and medium for monitoring the state of personnel in an elevator. Background Art

[0002] In an elevator with normal operation and a comfortable riding environment, the facial expressions and overall postures of passengers are usually relatively relaxed. If the elevator malfunctions, such as sudden stoppage of operation, shaking of the car, etc., or if a passenger suddenly falls ill, the mood of the passengers will immediately become anxious. They may worry about their safety and start looking around for an emergency call button or other rescue methods. At this time, their facial expressions and body postures will show significant changes. It is crucial to make a quick rescue response quickly and accurately based on the changes in the facial expressions and body postures of the personnel in the elevator.

[0003] However, the existing means for monitoring personnel in an elevator mainly include camera monitoring, sensor monitoring, etc., and have the following deficiencies: Most camera monitoring adopts the method of manual monitoring and cannot quickly and accurately judge the state of personnel in the elevator through the pictures captured by the camera; at the same time, sensor monitoring also cannot accurately judge the state of passengers in an environment where personnel in the elevator are relatively dense. Summary of the Invention

[0004] The main purpose of the present application is to provide a method, device, terminal and medium for monitoring the state of personnel in an elevator, aiming to accurately judge the state of personnel in the elevator by combining different angles through facial recognition and overall posture recognition, so that when sudden abnormal situations occur to the personnel in the elevator, the staff can handle them in a timely and rapid manner.

[0005] To achieve the above object, the present application provides a method for monitoring the state of personnel in an elevator, and the method includes:

[0006] Obtain the facial image and the overall image of the personnel in the elevator;

[0007] Through a preset first model, obtain a facial recognition score according to the facial image, and determine a first weight based on the facial recognition score, where the facial recognition score is used to represent the facial recognition state of the personnel in the elevator obtained through the preset first model;

[0008] Through a preset second model, obtain a posture recognition score according to the overall image, and determine a second weight based on the posture recognition score, where the posture recognition score is used to represent the body posture recognition state of the personnel in the elevator obtained through the preset second model;

[0009] Based on the face recognition score, the pose recognition score, the first weight, and the second weight, calculate the status score corresponding to the person in the elevator, and based on the status score, determine the target status monitoring result corresponding to the person in the elevator.

[0010] Specifically, the preset first model includes a first input layer, a first multi-scale feature extraction layer, a spatial pyramid pooling layer, a convolutional long short-term memory network layer, a global average pooling layer, a first fully connected layer, and a first output layer;

[0011] Obtaining the face recognition score according to the face image through the preset first model includes:

[0012] Through the first input layer, obtain a first image tensor according to the face image;

[0013] Through the first multi-scale feature extraction layer, obtain a set of multi-scale feature maps according to the first image tensor;

[0014] Through the spatial pyramid pooling layer, obtain a first matrix according to the set of multi-scale feature maps;

[0015] Through the convolutional long short-term memory network layer, obtain a first three-dimensional tensor according to the first matrix;

[0016] Through the global average pooling layer, obtain a first two-dimensional tensor according to the first three-dimensional tensor;

[0017] Through the first fully connected layer, obtain a first probability vector according to the first two-dimensional tensor;

[0018] Through the first output layer, obtain the face recognition score according to the first probability vector.

[0019] Specifically, determining the first weight based on the face recognition score includes:

[0020] According to the face recognition score, calculate the first weight through the following formula:

[0021]

[0022] Where, W f is the first weight, α and β are the first preset parameter and the second preset parameter respectively, and F is the face recognition score.

[0023] Specifically, the preset second model includes a second input layer, a key point detection layer, a graph construction layer, a graph convolution layer, a second multi-scale feature extraction layer, a global pooling layer, a second fully connected layer, and a second output layer;

[0024] Said obtaining the pose recognition score according to the whole body image through a preset second model includes:

[0025] Obtaining a second image tensor according to the whole body image through the second input layer;

[0026] Obtaining a key point coordinate matrix according to the second image tensor through the key point detection layer;

[0027] Obtaining an adjacency matrix and a feature matrix according to the key point coordinate matrix through the graph construction layer;

[0028] Obtaining a convolution feature matrix according to the adjacency matrix and the feature matrix through the graph convolution layer;

[0029] Obtaining a comprehensive feature matrix according to the convolution feature matrix through the two multi-scale feature extraction layer;

[0030] Obtaining a feature vector according to the comprehensive feature matrix through the global pooling layer;

[0031] Obtaining a second probability vector according to the feature vector through the second fully connected layer;

[0032] Obtaining the pose recognition score according to the second probability vector through the second output layer.

[0033] Specifically, said determining the second weight based on the pose recognition score includes:

[0034] Calculating the second weight according to the pose recognition score through the following formula:

[0035]

[0036] Wherein, W p is the second weight, tanh() represents the Tanh function, γ and δ are the third preset parameter and the fourth preset parameter respectively, and P is the pose recognition parameter.

[0037] Specifically, said calculating the state score corresponding to the person in the elevator based on the face recognition score, the pose recognition score, the first weight and the second weight includes:

[0038] Calculating the state score through the following formula:

[0039]

[0040] Wherein, S is the state score, W p is the second weight, P is the pose recognition parameter, W fis the first weight, and F is the face recognition score.

[0041] Specifically, the target status monitoring result is any one of a normal state, a warning state, and an abnormal state;

[0042] Determining the target status monitoring result corresponding to the person in the elevator based on the status score includes:

[0043] If the status score is within the first preset score range corresponding to the normal state, it is determined that the target status monitoring result is that the person in the elevator is in the normal state;

[0044] If the status score is within the second preset score range corresponding to the warning state, it is determined that the target status monitoring result is that the person in the elevator is in the warning state;

[0045] If the status score is within the third preset score range corresponding to the abnormal state, it is determined that the target status monitoring result is that the person in the elevator is in the abnormal state.

[0046] To achieve the above object, the present application further provides a device for monitoring the status of a person in an elevator. The device includes:

[0047] A first unit, configured to obtain target area information corresponding to a target area, sand mining operation area information corresponding to the sand mining operation area in the target area information, equipment information of a target ship, and navigation information of the target ship in the target area information, where the target area information is used to characterize the condition of the monitored water area, and the sand mining operation area is used to characterize the condition of the sand mining operation area in the target area;

[0048] A second unit, configured to obtain a first evaluation result according to the target area information, the sand mining operation area information, and the navigation information through a preset first model, where the first evaluation result is used to characterize the evaluation of the navigation condition of the target ship in the target area;

[0049] A third unit, configured to obtain a second evaluation result according to the equipment information and the pre-stored sand mining operation equipment information through a preset second model, where the second evaluation result is used to characterize the evaluation of the equipment condition on the target ship;

[0050] A fourth unit, configured to determine a target recognition result based on the first evaluation result, the second evaluation result, and a preset evaluation result threshold, where the target recognition result is used to characterize the determination result of whether the target ship implements sand mining operation in the target area.

[0051] To achieve the above object, the present application further provides a terminal, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any of the methods provided by the present application.

[0052] To achieve the above object, the present application further provides a medium storing multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any of the methods provided by the present application.

[0053] A method, device, terminal and medium for monitoring the state of personnel in an elevator provided by the present application may first obtain the facial image and full-body image of the personnel in the elevator; then, through a preset first model, based on the facial image, obtain a facial recognition score, and based on the facial recognition score, determine a first weight, where the facial recognition score is used to represent the facial recognition state of the personnel in the elevator obtained through the preset first model; then, through a preset second model, based on the full-body image, obtain a pose recognition score, and based on the pose recognition score, determine a second weight, where the pose recognition score is used to represent the body pose recognition state of the personnel in the elevator obtained through the preset second model; finally, based on the facial recognition score, the pose recognition score, the first weight and the second weight, calculate the state score corresponding to the personnel in the elevator, and based on the state score, accurately determine the target state monitoring result corresponding to the personnel in the elevator.

[0054] The present application can accurately judge the state of the personnel in the elevator from different angles through facial recognition and full-body pose recognition, so that when an unexpected abnormal situation occurs to the personnel in the elevator, the staff can handle it in a timely and rapid manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present application;

[0056] Figure 2 It is a schematic structural diagram of the device provided by the embodiment of the present application;

[0057] Figure 3 It is a schematic structural diagram of the terminal provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0059] Since the existing means for monitoring people in elevators mainly include camera monitoring, sensor monitoring, etc., there are the following deficiencies: Most camera monitoring uses manual viewing methods and cannot quickly and accurately judge the status of people in the elevator through the images captured by the cameras. At the same time, sensor monitoring also cannot accurately judge the status of passengers in an environment where people in the elevator are relatively dense.

[0060] Therefore, the embodiments of the present application provide a method, device, terminal, and medium for monitoring the status of people in an elevator to solve practical technical problems.

[0061] In some embodiments, the device may be specifically integrated in an electronic device, and the electronic device may be a device such as a terminal or a server.

[0062] In some embodiments, the server may also be implemented in the form of a terminal.

[0063] Among them, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0064] Among them, the terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.

[0065] The following will be described in detail respectively. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0066] The embodiments of the present application provide a method for monitoring the status of people in an elevator, as Figure 1 shown, the specific process of the method may be as follows:

[0067] S110. Obtain the facial image and full-body image of the people in the elevator.

[0068] In some embodiments, the facial image and full-body image of the people in the elevator may be obtained through the following hardware devices:

[0069] High-definition camera: Select a camera with a high resolution (such as 1080P and above) to ensure that the captured facial and full-body images are clear enough. For example, some professional surveillance cameras, like the DS-2CD2347G2-LU / SL model from Hikvision, can provide clear image details and meet the requirements for capturing the facial features and full-body contours of people.

[0070] Ultra-wide-angle camera: To obtain full-body images in a relatively narrow space like an elevator, an ultra-wide-angle lens is a good choice. Its field of view is relatively wide and can cover most areas from the top to the bottom and both sides of the elevator.

[0071] Camera with infrared function: If the light inside the elevator is dim or there may be light changes, a camera with an infrared function can ensure normal image acquisition in low-light or no-light environments. For example, at night or when the emergency lighting is insufficient during a power outage in the elevator, an infrared camera can use infrared light to form an image and still clearly capture the images of people.

[0072] In some embodiments, the camera can be installed at the top corner position of the elevator car, which can obtain a better top-down view angle and is beneficial for shooting full-body images of people. And according to the shape and size of the elevator, the angle of the camera can be adjusted to make the entire interior space of the elevator within the field of view of the camera as much as possible. In addition to the top installation, a camera can also be installed on the side of the elevator car. The side camera can supplement the shooting of the side facial images and body details of people, especially when people stand on the edge of the elevator car, it can better capture the images of these positions. The installation height is generally about 1.5 - 2 meters from the ground, and this height can better align with the facial area of people.

[0073] In some embodiments, the following software can be used in combination with the above hardware devices to obtain the facial images and full-body images:

[0074] A professional image acquisition software can be selected, which can communicate with the camera and control the parameter settings of the camera (such as resolution, frame rate, exposure, etc.). For example, some security surveillance software, like VMS (Video Management Software), can manage multiple cameras simultaneously and can obtain the image data transmitted by the camera regularly or in real-time. The image acquisition software needs to support the preliminary processing of images, such as adjusting the contrast, brightness, etc. of the images, to optimize the image quality and ensure that the facial and full-body details are clearer and more distinguishable.

[0075] A face recognition software can be adopted. The face recognition software usually uses deep learning algorithms, such as convolutional neural network (CNN). By training on a large number of face images, it can identify the facial features of different people. For example, the face recognition algorithm of Megvii Technology has a high accuracy rate and can accurately identify the identity of people in a complex elevator environment.

[0076] Meanwhile, a human pose recognition software can be adopted. The human pose recognition software can be used to analyze the full-body images of people, such as determining whether a person is standing, bending, or in other poses. This is very useful for monitoring in some special scenarios, such as whether there are abnormal behaviors in the elevator. It can also achieve pose recognition by analyzing the positions of human joint points.

[0077] S120: Through a preset first model, according to the facial image, obtain a face recognition score, and based on the face recognition score, determine a first weight, where the face recognition score is used to represent the face recognition status of the person in the elevator obtained through the preset first model.

[0078] In some embodiments, the preset first model includes a first input layer, a first multi-scale feature extraction layer, a spatial pyramid pooling layer, a convolutional long short-term memory network layer, a global average pooling layer, a first fully connected layer, and a first output layer.

[0079] The first multi-scale feature extraction layer can be a network layer based on a multi-scale feature extraction module (Multi-Scale Feature Extraction Module). The multi-scale feature extraction module uses multiple convolutional kernels of different scales (such as 3x3, 5x5, 7x7) to extract different levels of features of the face. After each convolutional layer, a ReLU activation function and a batch normalization layer are followed. The output of the multi-scale feature extraction module is a set of multi-scale feature maps, and the size of each feature map is different according to the size of the convolutional kernel.

[0080] The spatial pyramid pooling layer can be a network layer based on spatial pyramid pooling (SPP). The spatial pyramid pooling layer can convert the multi-scale feature map into a fixed-length feature vector through SPP to adapt to inputs of different sizes. SPP divides the feature map into multiple regions and aggregates their features by performing max pooling operations at different scales.

[0081] The convolutional long short-term memory network layer can be a network layer based on the Convolutional Long Short-Term Memory (ConvLSTM). The convolutional long short-term memory network is used to capture the temporal dependence and dynamic changes of facial images and is particularly useful for continuous-frame face recognition tasks. The convolutional long short-term memory network contains multiple units, and each unit has a forget gate, an input gate, a cell state, and an output gate inside to control the flow of information.

[0082] The global average pooling layer can be a network layer based on Global Average Pooling (GAP). The global average pooling layer can perform global average pooling on the feature vectors at each time step to obtain a feature vector with a fixed length.

[0083] Specifically, obtaining the face recognition score according to the facial image through the preset first model includes the step contents from S121 to S127 as shown below:

[0084] S121: Through the first input layer, obtain a first image tensor according to the facial image.

[0085] In some embodiments, the facial image can be an RGB image of 224x224x3. Inputting the RGB image of 224x224x3 into the first input layer can output the first image tensor, and the first image tensor can be an image tensor of 224x224x3.

[0086] S122: Through the first multi-scale feature extraction layer, obtain a set of multi-scale feature maps according to the first image tensor.

[0087] Continuing the above embodiment for further explanation, inputting the image tensor of 224x224x3 into the first multi-scale feature extraction layer to obtain the set of multi-scale feature maps, and the set of multi-scale feature maps can be [224x224x64, 112x112x128, 56x56x256].

[0088] S123: Through the spatial pyramid pooling layer, obtain a first matrix according to the set of multi-scale feature maps.

[0089] Continuing the above embodiment for further explanation, inputting the set of multi-scale feature maps into the spatial pyramid pooling layer and outputting the first matrix. The first matrix can be a matrix of [N, D], where N = 21 (such as 1x1, 2x2, 4x4 pooling), and D = 256.

[0090] S124. Obtain a first three-dimensional tensor according to the first matrix through the convolutional long short-term memory network layer.

[0091] Continuing the description based on the above embodiment, input the matrix of [N, D] into the convolutional long short-term memory network layer, and output the first three-dimensional tensor. The first three-dimensional tensor can be a three-dimensional tensor of [T, N, D], where T is the number of frames.

[0092] S125. Obtain a first two-dimensional tensor according to the first three-dimensional tensor through the global average pooling layer.

[0093] Continuing the description based on the above embodiment, input the three-dimensional tensor of [T, N, D] into the global average pooling layer, and output the first two-dimensional tensor. The first two-dimensional tensor can be a two-dimensional tensor of [T, D].

[0094] S126. Obtain a first probability vector according to the first two-dimensional tensor through the first fully connected layer.

[0095] Continuing the description based on the above embodiment, input the two-dimensional tensor of [T, D] into the first fully connected layer, and obtain the first probability vector which can be a probability vector of [C1].

[0096] S127. Obtain the face recognition score according to the first probability vector through the first output layer.

[0097] Continuing the description based on the above embodiment, input the probability vector of [C1] into the first output layer, and output the face recognition score.

[0098] In some embodiments, determining the first weight based on the face recognition score includes the following specific implementation process:

[0099] Calculate the first weight according to the face recognition score through the following formula:

[0100]

[0101] where, W f is the first weight, β and α are the first preset parameter and the second preset parameter respectively, and F is the face recognition score.

[0102] S130. Obtain a pose recognition score according to the full body image through a preset second model, and determine a second weight based on the pose recognition score, where the pose recognition score is used to represent the body pose recognition state of the person in the elevator obtained through the preset second model.

[0103] In some embodiments, the preset second model includes a second input layer, a keypoint detection layer, a graph construction layer, a graph convolutional layer, a second multi-scale feature extraction layer, a global pooling layer, a second fully connected layer, and a second output layer.

[0104] The keypoint detection layer can be a network layer based on Keypoint Detection, and the human body keypoints (such as head, shoulders, elbows, wrists, hips, knees, ankles, etc.) in the image can be detected through the keypoint detection layer (such as OpenPose, HRNet, etc.).

[0105] The graph construction layer can construct an undirected graph according to the detected keypoints, where each keypoint is a node, and the edges represent the connection relationships between the keypoints (such as limbs, torso, etc.).

[0106] The graph convolutional layer is used to perform convolutional operations on the graph to capture the spatial relationships between nodes, and different aggregation functions (such as average pooling, max pooling, etc.) can be used to aggregate the information of neighboring nodes.

[0107] The second multi-scale feature extraction layer can be a network layer based on Multi-Scale Feature Extraction. Multiple graph convolutional layers with different scales can be used to gradually extract features from local to global, and finally the feature matrices of all scales are concatenated together to form a comprehensive feature matrix.

[0108] The global pooling layer can perform global pooling operations (such as max pooling or average pooling) on the comprehensive feature matrix to obtain a feature vector with a fixed length.

[0109] Specifically, obtaining the pose recognition score according to the full body image through the preset second model includes the steps S131 to S138 as shown below:

[0110] S131. Through the second input layer, obtain a second image tensor according to the full body image.

[0111] In some embodiments, the full body image can be an RGB image of 368x368x3, and the second image tensor can be an image tensor of 368x368x3.

[0112] S132. Through the keypoint detection layer, obtain a keypoint coordinate matrix according to the second image tensor.

[0113] Continuing the description based on the above embodiments, the image tensor of 368x368x3 is input into the keypoint detection layer, and the output can be a keypoint coordinate matrix of [K, 2], where K is the number of keypoints.

[0114] S133. Through the graph construction layer, an adjacency matrix and a feature matrix are obtained according to the key point coordinate matrix.

[0115] Continuing the description of the above embodiment, the key point coordinate matrix of [K, 2] is input into the graph construction layer, and an adjacency matrix A ∈ R can be output. K×K and a feature matrix X ∈ R K×L , where L is the feature dimension of each key point (such as coordinates, confidence, etc.).

[0116] S134. Through the graph convolution layer, a convolution feature matrix is obtained according to the adjacency matrix and the feature matrix.

[0117] Continuing the description of the above embodiment, the adjacency matrix and the feature matrix are input into the graph convolution layer, and a convolution feature matrix [K, M] can be output, where M is the feature dimension of the graph convolution layer.

[0118] S135. Through the two - multi - scale feature extraction layer, a comprehensive feature matrix is obtained according to the convolution feature matrix.

[0119] Continuing the description of the above embodiment, the convolution feature matrix [K, M] is input into the two - multi - scale feature extraction layer, and a comprehensive feature matrix [K, ∑M i is output.

[0120] S136. Through the global pooling layer, a feature vector is obtained according to the comprehensive feature matrix.

[0121] Continuing the description of the above embodiment, the comprehensive feature matrix [K, ∑M i is input into the global pooling layer, and a feature vector [M'] is output.

[0122] S137. Through the second fully - connected layer, a second probability vector is obtained according to the feature vector.

[0123] Continuing the description of the above embodiment, the feature vector [M'] is input into the second fully - connected layer, and a second probability vector [C2] is output.

[0124] S138. Through the second output layer, the pose recognition score is obtained according to the second probability vector.

[0125] Continuing the description of the above embodiment, the second probability vector [C2] is input into the second output layer, and the pose recognition score is output.

[0126] In some embodiments, determining the second weight based on the pose recognition score includes the following specific implementation process:

[0127] According to the pose recognition score, the second weight is calculated through the following formula:

[0128]

[0129] where W p is the second weight, tanh() represents the Tanh function, γ and δ are the third preset parameter and the fourth preset parameter respectively, and P is the pose recognition parameter.

[0130] S140. Based on the face recognition score, the pose recognition score, the first weight, and the second weight, calculate the status score corresponding to the person in the elevator, and based on the status score, determine the target status monitoring result corresponding to the person in the elevator.

[0131] In some embodiments, the calculating the status score corresponding to the person in the elevator based on the face recognition score, the pose recognition score, the first weight, and the second weight includes the following specific implementation process:

[0132] The status score is calculated through the following formula:

[0133]

[0134] where S is the status score, W p is the second weight, P is the pose recognition parameter, W f is the first weight, and F is the face recognition score.

[0135] In some embodiments, the target status monitoring result is any one of a normal state, a warning state, and an abnormal state.

[0136] Specifically, the determining the target status monitoring result corresponding to the person in the elevator based on the status score includes the following specific implementation process:

[0137] If the status score is within the first preset score range corresponding to the normal state, it is determined that the target status monitoring result is that the person in the elevator is in the normal state;

[0138] If the status score is within the second preset score range corresponding to the warning state, it is determined that the target status monitoring result is that the person in the elevator is in the warning state;

[0139] If the status score is within the third preset score range corresponding to the abnormal state, it is determined that the target status monitoring result is that the person in the elevator is in the abnormal state.

[0140] Specifically, the status score S of the people in the elevator has been calculated, and it ranges from [0, 1]. A simple classification rule can be designed to determine the target status monitoring result. Let R be the target status monitoring result, which can take values of "Normal", "Warning", or "Abnormal", used to represent the normal state, warning state, or abnormal state.

[0141] The first preset score interval can be [0.8, 1]. When the status score S is not less than 0.8, the people in the elevator are in a normal state.

[0142] The second preset score interval can be [0.5, 0.8). When the status score S is not less than 0.5 and less than 0.8, the people in the elevator are in a warning state. When in the warning state, it may mean that the people in the elevator are very likely to be in a sudden emergency situation, which is worthy of the attention of the monitoring personnel, or the emergency rescue system is activated and on standby at any time.

[0143] The third preset score interval can be [0, 0.5). When the status score S is not less than 0 and less than 0.5, the people in the elevator are in an abnormal state, and the emergency rescue system is triggered, and the rescue personnel go to deal with it in time.

[0144] In summary, the present application provides a method for monitoring the status of people in an elevator. By identifying the navigation information of the target ship and the equipment on the ship, the target ship is comprehensively evaluated from multiple angles, and then it is accurately determined whether the target ship is engaged in sand mining operations.

[0145] To better implement the above method, the embodiment of the present application also provides a device for monitoring the status of people in an elevator. This device can be specifically integrated in an electronic device, and the electronic device can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers.

[0146] For example, in this embodiment, taking the device for monitoring the status of people in an elevator being specifically integrated in a terminal as an example, the method of the embodiment of the present application will be described in detail.

[0147] For example, as Figure 2 shown, the device 200 for monitoring the status of people in an elevator can include a first unit 201, a second unit 202, a third unit 203, and a fourth unit 204. The device includes:

[0148] The first unit is used to obtain the facial image and the full-body image of the people in the elevator;

[0149] A second unit is configured to obtain a face recognition score based on the face image through a preset first model, and determine a first weight based on the face recognition score, where the face recognition score is used to represent the face recognition status of the person in the elevator obtained through the preset first model;

[0150] A third unit is configured to obtain a posture recognition score based on the whole body image through a preset second model, and determine a second weight based on the posture recognition score, where the posture recognition score is used to represent the body posture recognition status of the person in the elevator obtained through the preset second model;

[0151] A fourth unit is configured to calculate a status score corresponding to the person in the elevator based on the face recognition score, the posture recognition score, the first weight, and the second weight, and determine a target status monitoring result corresponding to the person in the elevator based on the status score.

[0152] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0153] As can be seen from the above, the embodiments of the present application can accurately judge the status of the person in the elevator from different angles through face recognition and whole body posture recognition, so that when an unexpected abnormal situation occurs to the person in the elevator, the staff can handle it in a timely and rapid manner.

[0154] The embodiments of the present application further provide an electronic device, which can be a device such as a terminal or a server. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.

[0155] In some embodiments, the product processing device can also be integrated in multiple electronic devices. For example, the product processing device can be integrated in multiple servers, and the method for monitoring the status of the person in the elevator of the present application can be implemented by multiple servers.

[0156] In this embodiment, the electronic device in this embodiment is taken as an example of a terminal for detailed description. For example, as Figure 3 shown, it shows a schematic structural diagram of a terminal 300 involved in the embodiments of the present application. Specifically:

[0157] The terminal 300 may include a processor 301 with one or more processing cores, a memory 302 with one or more media, a power supply 303, an input module 304, and a communication module 305 and other components. Those skilled in the art can understand,Figure 3 The structure of the terminal 300 shown does not constitute a limitation on the terminal 300, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0158] The processor 301 is the control center of the terminal 300, connecting various parts of the entire terminal 300 through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 302, and calling the data stored in the memory 302, it executes various functions of the terminal 300 and processes data, thereby monitoring the terminal 300 as a whole. In some embodiments, the processor 301 may include one or more processing cores; in some embodiments, the processor 301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 301.

[0159] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created according to the use of the terminal 300. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0160] The terminal 300 also includes a power supply 303 for supplying power to each component. In some embodiments, the power supply 303 may be logically connected to the processor 301 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 303 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0161] The terminal 300 may also include an input module 304, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0162] The terminal 300 may further include a communication module 305. In some embodiments, the communication module 305 may include a wireless module. The terminal 300 may perform short-range wireless transmission through the wireless module of the communication module 305, thereby providing users with wireless broadband Internet access. For example, the communication module 305 may be used to help users send and receive emails, browse web pages, and access streaming media, etc.

[0163] Although not shown, the terminal 300 may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 301 in the terminal 300 will load the executable files corresponding to the processes of one or more application programs into the memory 302 according to the following instructions, and the processor 301 will run the application programs stored in the memory 302 to implement various functions as follows:

[0164] Obtain the facial image and full-body image of the person in the elevator;

[0165] Through a preset first model, according to the facial image, obtain a facial recognition score, and based on the facial recognition score, determine a first weight, where the facial recognition score is used to represent the facial recognition state of the person in the elevator obtained through the preset first model;

[0166] Through a preset second model, according to the full-body image, obtain a pose recognition score, and based on the pose recognition score, determine a second weight, where the pose recognition score is used to represent the body pose recognition state of the person in the elevator obtained through the preset second model;

[0167] Based on the facial recognition score, the pose recognition score, the first weight, and the second weight, calculate the state score corresponding to the person in the elevator, and based on the state score, determine the target state monitoring result corresponding to the person in the elevator.

[0168] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0169] As can be seen from the above, the embodiments of the present application can accurately judge the state of the person in the elevator from different angles through facial recognition and full-body pose recognition, so that when an unexpected abnormal situation occurs to the person in the elevator, the staff can handle it in a timely and rapid manner.

[0170] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a medium and loaded and executed by a processor.

[0171] To this end, an embodiment of the present application provides a medium in which multiple instructions are stored, and these instructions can be loaded by a processor to execute the steps in any one of the elevator passenger status monitoring methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:

[0172] Obtain the facial image and full-body image of the passengers in the elevator;

[0173] Through a preset first model, obtain a facial recognition score based on the facial image, and determine a first weight based on the facial recognition score, where the facial recognition score is used to represent the facial recognition status of the passengers in the elevator obtained through the preset first model;

[0174] Through a preset second model, obtain a pose recognition score based on the full-body image, and determine a second weight based on the pose recognition score, where the pose recognition score is used to represent the body pose recognition status of the passengers in the elevator obtained through the preset second model;

[0175] Based on the facial recognition score, the pose recognition score, the first weight, and the second weight, calculate the status score corresponding to the passengers in the elevator, and determine the target status monitoring result corresponding to the passengers in the elevator based on the status score.

[0176] Among them, the medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0177] According to an aspect of the present application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions, and the computer instructions are stored in a medium. The processor of the computer device reads the computer instructions from the medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementation manners provided in the above embodiments.

[0178] Since the instructions stored in the medium can execute the steps in any one of the elevator passenger status monitoring methods provided by the embodiments of the present application, the beneficial effects that can be achieved by any one of the elevator passenger status monitoring methods provided by the embodiments of the present application can be realized. For details, see the previous embodiments and will not be elaborated here.

[0179] The above has introduced in detail a method, device, terminal, and medium for monitoring the status of personnel in an elevator provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for monitoring the status of personnel in an elevator, characterized in that, The method includes: Obtaining a facial image and a full-body image of a person in the elevator; Through a preset first model, according to the facial image, obtaining a facial recognition score, and based on the facial recognition score, determining a first weight, where the facial recognition score is used to characterize the facial recognition status of the person in the elevator obtained through the preset first model; Through a preset second model, according to the full-body image, obtaining a pose recognition score, and based on the pose recognition score, determining a second weight, where the pose recognition score is used to characterize the body pose recognition status of the person in the elevator obtained through the preset second model; Based on the facial recognition score, the pose recognition score, the first weight, and the second weight, calculating to obtain a state score corresponding to the person in the elevator, and based on the state score, determining a target state monitoring result corresponding to the person in the elevator.

2. The method according to claim 1, characterized in that The preset first model includes a first input layer, a first multi-scale feature extraction layer, a spatial pyramid pooling layer, a convolutional long short-term memory network layer, a global average pooling layer, a first fully connected layer, and a first output layer; The obtaining the facial recognition score according to the facial image through the preset first model includes: Through the first input layer, according to the facial image, obtaining a first image tensor; Through the first multi-scale feature extraction layer, according to the first image tensor, obtaining a set of multi-scale feature maps; Through the spatial pyramid pooling layer, according to the set of multi-scale feature maps, obtaining a first matrix; Through the convolutional long short-term memory network layer, according to the first matrix, obtaining a first three-dimensional tensor; Through the global average pooling layer, according to the first three-dimensional tensor, obtaining a first two-dimensional tensor; Through the first fully connected layer, according to the first two-dimensional tensor, obtaining a first probability vector; Through the first output layer, according to the first probability vector, obtaining the facial recognition score.

3. The method according to claim 1, wherein The determining the first weight based on the facial recognition score includes: According to the facial recognition score, calculating to obtain the first weight through the following formula: Among them, W f is the first weight, α and β are the first preset parameter and the second preset parameter respectively, and F is the face recognition score.

4. The method according to claim 1, characterized in that, The preset second model includes a second input layer, a key point detection layer, a graph construction layer, a graph convolutional layer, a second multi-scale feature extraction layer, a global pooling layer, a second fully connected layer, and a second output layer; The obtaining the pose recognition score according to the full-body image through the preset second model includes: Through the second input layer, according to the full-body image, obtaining a second image tensor; Through the key point detection layer, according to the second image tensor, obtaining a key point coordinate matrix; Through the graph construction layer, according to the key point coordinate matrix, obtaining an adjacency matrix and a feature matrix; Through the graph convolutional layer, according to the adjacency matrix and the feature matrix, obtaining a convolutional feature matrix; Through the second multi-scale feature extraction layer, according to the convolutional feature matrix, obtaining a comprehensive feature matrix; Through the global pooling layer, according to the comprehensive feature matrix, obtaining a feature vector; Through the second fully connected layer, according to the feature vector, obtaining a second probability vector; Through the second output layer, obtain the pose recognition score according to the second probability vector.

5. The method according to claim 1, wherein Determining the second weight based on the pose recognition score includes: According to the pose recognition score, calculate the second weight through the following formula: Among them, W p is the second weight, tanh() represents the Tanh function, γ and δ are the third preset parameter and the fourth preset parameter respectively, and P is the pose recognition parameter.

6. The method according to claim 1, wherein Calculating the state score corresponding to the person in the elevator based on the face recognition score, the pose recognition score, the first weight, and the second weight includes: Calculate the state score through the following formula: Among them, S is the state score, and W p is the second weight, P is the pose recognition parameter, and W f is the first weight, and F is the face recognition score.

7. The method according to claim 1, wherein The target state monitoring result is any one of a normal state, a warning state, and an abnormal state; Determining the target state monitoring result corresponding to the person in the elevator based on the state score includes: If the state score is within the first preset score range corresponding to the normal state, determine that the target state monitoring result is that the person in the elevator is in the normal state; If the state score is within the second preset score range corresponding to the warning state, determine that the target state monitoring result is that the person in the elevator is in the warning state; If the state score is within the third preset score range corresponding to the abnormal state, determine that the target state monitoring result is that the person in the elevator is in the abnormal state.

8. An in-elevator personnel status monitoring device, characterized in that, The device includes: A first unit for acquiring the face image and the full body image of the person in the elevator; A second unit for obtaining a face recognition score according to the face image through a preset first model, and determining a first weight based on the face recognition score, where the face recognition score is used to represent the face recognition state of the person in the elevator obtained through the preset first model; A third unit for obtaining a pose recognition score according to the full body image through a preset second model, and determining a second weight based on the pose recognition score, where the pose recognition score is used to represent the body pose recognition state of the person in the elevator obtained through the preset second model; A fourth unit for calculating the state score corresponding to the person in the elevator based on the face recognition score, the pose recognition score, the first weight, and the second weight, and determining the target state monitoring result corresponding to the person in the elevator based on the state score.

9. A terminal, characterized in that, It includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in the method according to any one of claims 1 to 7.

10. A medium, characterized in that, The medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the method according to any one of claims 1 to 7.