Deep Learning-Based Passenger Multi-Label Classification Method, System, Medium and Device

Through the passenger multi-label classification system based on deep learning, the problem of multi-label attribute identification for passengers in the elevator is solved, and the functions of the elevator intelligent monitoring system are improved and the safety and commerciality of elevator riding are improved.

CN114973325BActive Publication Date: 2025-06-10SHANGHAI JIAOTONG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202210555443.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-06-10
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and classify the multi-label attributes of passengers in elevators, which limits the functions and applications of the elevator intelligent monitoring system.

Method used

A passenger multi-label classification system based on deep learning is adopted to realize intelligent identification and classification of passenger multi-label attributes through modules such as image data acquisition, attribute data set establishment and processing, and passenger attribute recognition.

Benefits of technology

It realizes the accurate identification and classification of multi-label attributes of passengers in the elevator, improves the functions and applications of the elevator intelligent monitoring system, and improves the safety and commerciality of elevator riding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973325B_ABST
    Figure CN114973325B_ABST
Patent Text Reader

Abstract

The present invention provides a passenger multi-label classification system, method, medium and device based on deep learning, including: an image data acquisition module: acquiring a surveillance video, transcoding and playing it, calling a target detection algorithm to perform personnel detection and intercepting passenger images; an attribute data set establishment and processing module: establishing a passenger attribute data set, traversing and proofreading images and labels, respectively extracting image and label lists, and converting the labels into one-hot codes; a passenger attribute recognition module: loading images and preprocessing them, loading attribute labels and rearranging them, dividing the data set, configuring deep neural network model parameters and training, performing model prediction, and calculating evaluation indicators at the image level and the attribute level; a result display and saving module: constructing a background and prediction dictionary, printing the attribute recognition and evaluation results and saving them. The present invention can achieve the recognition of key passengers, provide technical support for work such as elevator big data, intelligent security, and advertising placement, and improve the safety and commerciality of elevator rides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a multi-label classification method, system, medium and device for passengers based on deep learning. Background Art

[0002] Elevators are an essential means of transportation in high-rise buildings. In the era of rapid economic and technological development, urbanization and artificial intelligence have been widely popularized in daily life. Traditional monitoring systems have a single function and manual inspections are laborious and time-consuming. With the sharp increase in the use of surveillance cameras and the development of machine learning, image-based crowd scene analysis has attracted great interest in the academic and industrial fields in recent years. Building an intelligent elevator monitoring system has gradually become one of the focus research directions in machine vision.

[0003] Patent document CN111144247A (application number: CN201911292323.1) discloses a method for detecting the retrograde behavior of escalator passengers based on deep learning, including the steps of: Step 1, obtaining image frames from the surveillance video stream of the escalator and setting a detection region (ROI); Step 2, using an object detection algorithm to detect the head position of passengers in the ROI region specified in Step 1; Step 3, using a classifier to discriminate the direction of the human head detected in Step 2 to determine whether the passenger is likely to exhibit a retrograde phenomenon; Step 4, using a multi-object tracking algorithm to track each possible retrograde target (head) in Step 3 to obtain a tracking trajectory; Step 5, analyzing each trajectory in Step 4 to determine whether the passenger has committed a retrograde behavior.

[0004] Patent document CN108805093A (application number: CN201810627161.1) discloses a detection algorithm for detecting the fall of escalator passengers based on deep learning, including the steps of: 1) collecting video images of passengers on the escalator; 2) detecting the faces of passengers using the FHOG descriptor and the SVM classifier; 3) tracking the faces of passengers using KCF and creating a new passenger trajectory list based on the face information of the passengers; 4) retraining the yolo2 algorithm model using transfer learning to detect the human body of the passengers; 5) matching the faces of the passengers and the human body of the passengers and adding the body information to the trajectory list; 6) using the openpose deep learning algorithm to extract the sequence of skeletal joint points of the passengers; 7) matching the human body of the passengers and the sequence of skeletal joint points of the passengers and adding the skeletal joint point information to the trajectory list; 8) analyzing the skeletal joint point information in the trajectory list to detect the fall behavior of the passengers.

[0005] Patent document CN111259718A (application number: CN201911017881.7) discloses an escalator detention detection method and system based on a Gaussian mixture model, including the following steps: A) using a deep learning network to extract pedestrian information from video images; B) using a machine learning method to build a classification model to classify and screen the pedestrian information detected in real time; C) building a Gaussian mixture model to obtain the detention target image of the current frame; D) preprocessing the detention target image to obtain an image within the detention detection area at the escalator entrance and exit; E) using an edge detection algorithm to obtain the outline of the candidate detention target and perform candidate detention box area judgment; F) performing candidate detention box overlap judgment and performing a detention alarm.

[0006] Crowd attribute recognition belongs to the eye category of fine-grained visual classification. Fine-grained image analysis, as an interesting, basic and challenging problem in computer vision, has been an active research field for decades. Its goal is to retrieve, identify and generate images belonging to multiple subordinate categories under a supercategory. It can be divided into two categories: strong supervision-based methods and weak supervision-based methods. The former relies on more manual annotations (image labels, target locations, etc.) to learn to detect targets, and has high requirements for data construction. The weak supervision method allows the model to focus on the local part of the image, so that the update of network parameters is carried out towards mining more distinguishable image areas. In recent years, the rise of deep learning has led researchers to mostly use deep convolutional neural networks to build pedestrian attribute recognition models. According to the focus of model establishment, the current numerous pedestrian attribute recognition methods can be classified into five categories: image block-based methods, whole image-based methods, attention-based methods, serialized feature learning-based methods, and weak supervision-based methods. Most of the current pedestrian attribute recognition methods involve fine-tuning of model parameters in transfer learning to make up for the problem of insufficient training data.

[0007] Elevator cars are small, closed, and opaque. Intelligent monitoring aims to use deep learning, video processing and other methods to analyze monitoring videos, dynamically monitor passengers, and mine the data patterns behind them. This can create conditions for intelligent management of elevators and accurate placement of elevator advertisements, improve the safety and commerciality of elevator rides, and help change traditional elevator management and operation models. Therefore, the research on intelligent classification methods of multi-label attributes of elevator passengers based on deep learning has important research value and significance for building intelligent monitoring of elevator cars. Summary of the invention

[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a passenger multi-label classification method, system, medium and equipment based on deep learning.

[0009] The multi-label classification system for passengers based on deep learning provided by the present invention includes:

[0010] An image data acquisition module: acquiring a surveillance video, transcoding and playing it, and calling a target detection algorithm to detect people and intercept passenger images;

[0011] An attribute data set establishment and processing module: establishing a passenger attribute data set, traversing and proofreading images and labels, respectively extracting image and label lists, and converting the labels into one-hot codes;

[0012] A passenger attribute recognition module: first loading and preprocessing images, loading attribute labels and rearranging them, then dividing the data set, configuring deep neural network model parameters and training, and then performing model prediction, and calculating evaluation metrics at the image level and the attribute level;

[0013] A result display and saving module: constructing a background and prediction dictionary, printing the attribute recognition and evaluation results and saving them.

[0014] Preferably, the image data acquisition module includes a surveillance video acquisition unit, a person detection unit, and an image interception unit that are communicatively connected in sequence;

[0015] The surveillance video acquisition unit uses a camera installed on the top of the elevator to collect surveillance videos in real time, transcode them, and play them smoothly;

[0016] The person detection unit calls a deep learning method to detect people in the video stream;

[0017] The image interception unit is used to intercept the person detection regression box in the video stream and save it to the data set path.

[0018] Preferably, the attribute data set establishment and processing module includes:

[0019] Establishing an elevator image data set, including an image folder and a label csv file;

[0020] Using the method for establishing and processing the passenger attribute data set to traverse and proofread the images and labels to make them correspond one by one;

[0021] Respectively extracting image and label lists, and converting the text labels into binary one-hot codes.

[0022] Preferably, the passenger attribute recognition module includes:

[0023] A model construction and training unit for preparing training data, configuring network parameters, and training a deep learning model;

[0024] A model testing and evaluation unit for evaluating the model performance at the image level and the attribute level respectively, and the evaluation metrics include accuracy, precision, and recall.

[0025] A multi-label classification method for passengers based on deep learning provided by the present invention includes:

[0026] An image data acquisition step: obtaining a surveillance video, transcoding and playing it, and calling a target detection algorithm to detect people and intercept passenger images;

[0027] An attribute dataset establishment and processing step: establishing a passenger attribute dataset, traversing and proofreading images and labels, respectively extracting image and label lists, and converting the labels into one-hot codes;

[0028] A passenger attribute recognition step: first loading an image and preprocessing it, loading attribute labels and rearranging them, then dividing the dataset, configuring deep neural network model parameters and training, and then performing model prediction, and calculating evaluation metrics at the image level and the attribute level;

[0029] A result display and saving step: constructing a background and prediction dictionary, printing the attribute recognition and evaluation results, and saving them.

[0030] Preferably, the image data acquisition step includes:

[0031] A surveillance video acquisition step: using a camera installed on the top of the elevator to collect a surveillance video in real time, transcoding it, and playing it smoothly;

[0032] A personnel detection step: calling a deep learning method to detect people in the video stream;

[0033] An image interception step: intercepting the personnel detection regression box in the video stream and saving it to the dataset path.

[0034] Preferably, the attribute dataset establishment and processing step includes:

[0035] Establishing an elevator image dataset, including an image folder and a label csv file;

[0036] Using a method for establishing and processing a passenger attribute dataset to traverse and proofread images and labels to make them correspond one by one;

[0037] Respectively extracting image and label lists, and converting the text labels into binary one-hot codes.

[0038] Preferably, the passenger attribute recognition step includes:

[0039] A model construction and training step: preparing training data, configuring network parameters, and training a deep learning model;

[0040] A model testing and evaluation step: evaluating the model performance at the image level and the attribute level respectively, and the evaluation metrics include accuracy, precision, and recall.

[0041] A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the method are implemented.

[0042] A deep learning-based multi-label classification device for passengers according to the present invention includes: a controller; the controller includes the computer-readable storage medium storing the computer program, and when the computer program is executed by a processor, the steps of the deep learning-based multi-label classification method for passengers are implemented; alternatively, the controller includes the deep learning-based multi-label classification system for passengers.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention can realize the identification of key passengers, provide technical support for work such as elevator big data, intelligent security, and advertising placement, and improve the safety and commerciality of elevator rides. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:

[0046] Figure 1 It is a schematic flowchart of a deep learning-based intelligent multi-label attribute classification method for elevator passengers provided by the present invention;

[0047] Figure 2 It is a schematic composition diagram of an elevator image data acquisition device provided by the present invention;

[0048] Figure 3 It is a schematic flowchart of a method for establishing and processing an elevator passenger attribute data set provided by the present invention;

[0049] Figure 4 It is a schematic flowchart of model construction and training in the intelligent multi-label attribute recognition method for elevator passengers provided by the present invention;

[0050] Figure 5 It is a schematic flowchart of model testing and evaluation in the intelligent multi-label attribute recognition method for elevator passengers provided by the present invention;

[0051] Figure 6 It is a schematic structural diagram of an elevator passenger attribute recognition device provided by the present invention;

[0052] Figure 7 It is a schematic structural diagram of a computer device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several changes and improvements can still be made. These all belong to the protection scope of the present invention.

[0054] Embodiment 1:

[0055] First, the present invention provides an elevator image data acquisition device, including a monitoring video acquisition unit, a personnel detection unit, and an image interception unit that are communicatively connected in sequence.

[0056] The monitoring video acquisition unit uses a camera installed on the top of the elevator to collect monitoring videos in real time, transcode them, and play them smoothly.

[0057] The personnel detection unit calls deep learning methods to detect personnel in the video stream. Optional target detection algorithms include: single-stage detectors YOLO, SSD, etc. that directly perform classification and regression based on feature extraction, and two-stage detectors Fast R-CNN, etc. based on candidate regions.

[0058] The image interception unit is used to intercept the personnel detection regression box in the elevator video stream and save it to the dataset path.

[0059] Second, the present invention provides a method for establishing and processing an elevator passenger attribute dataset, including the following steps:

[0060] 1. Obtain an elevator image dataset, including an image folder and a label csv file.

[0061] 2. Process the dataset, traverse and proofread the images and labels to make them correspond one by one.

[0062] 3. Extract the image img and label list respectively.

[0063] 4. Convert the label to a binary one-hot code.

[0064] 5. Save the binarized multi-label as a pkl file.

[0065] In step 1, obtain the elevator passenger images and attribute labels collected in the first aspect. The image names are named with increasing Arabic numerals in the order of interception. The first column of the csv file is the image name, and the 2nd - 12th columns are the corresponding label information.

[0066] In the said step 1, the label attributes include the following categories: gender (male, female); age (child, teenager, adult, elderly); headgear (without hat, with hat, with helmet); hair length (unidentifiable, bald, short hair, long hair); whether white hair (non-white hair, white hair); whether glasses (without glasses, with glasses); occupation (courier, cleaner, resident); whether holding a child (not holding a child, holding a child); whether playing with mobile phone (not playing with mobile phone, playing with mobile phone); whether carrying a backpack (not carrying a backpack, carrying a backpack); whether holding a handbag (holding a handbag, holding a handbag), a total of 11 types.

[0067] In the said step 2, the specific steps include: ① Read the label file and the image folder respectively, that is, save all images and corresponding labels as list_1 and list_2 respectively; ② Compare the intersection to obtain the missing files. That is, first obtain the intersection list inter of list_1 and list_2, and then obtain the difference sets diff_1 between the image list list_1 and the intersection list inter, and the difference set diff_2 between the label list list_2 and the intersection list inter respectively. ③ Delete the incomplete difference set files. First, delete the images that do not exist in the csv file in the image folder, that is, delete diff_1 in list_1, and then delete the images that do not exist in the image folder in the csv file, that is, delete diff_2 in list_2. ④ Delete the duplicate rows in the label file. When the number of labels is greater than the number of images in the folder, but the difference set between the two is 0, it means that there are duplicate rows in the labels. After the above operations, the data set processing is completed, so that the images and labels correspond one by one.

[0068] In the said step 3, the row index of the csv file is the image name, and the column index is the attribute name. ① Read the first column of the csv file as a list, denoted as img, and img is in the form of [1.jpg, 2.jpg,... n.jpg], and the length is the number of images; ② Read the row corresponding to each img index as a list, denoted as label, and label is a text label with a length equal to 11.

[0069] In the said step 4, perform binary processing on the label list obtained in step 3. Specifically, let the attribute be recorded as 1 and the non-existence be recorded as 0, and convert label into a one-hot code. Then label is in the form of [101... 1], and the length is 27, which respectively represent the existence of the following 27 attributes: "not wearing a hat, not wearing glasses, not holding a child, not holding a handbag, not playing with mobile phone, not carrying a backpack, cleaner, bald / balding, hair unidentifiable, female, child, resident, courier, adult, wearing a helmet, wearing a hat, wearing glasses, holding a handbag, playing with mobile phone, male, white hair, short hair, elderly, carrying a backpack, long hair, teenager, non-white hair".

[0070] In step 5, using the binarization result of step 4, save the multi-label attribute classes as a pkl file for facilitating the test and printing of attribute information, where the length of the attribute classes is 27.

[0071] In a third aspect, the present invention provides an intelligent recognition method for multi-label attributes of elevator passengers, including a model construction and training unit and a model test and evaluation unit. Among them, the model construction and training include three steps: data preparation, model configuration, and model training and saving; the model test is divided into two parts: the image level and the attribute level, and the evaluation indicators include accuracy, precision, and recall.

[0072] (1) Model construction and training unit

[0073] 1. Data preparation:

[0074] (1) Configure the dataset and the path of the pre-trained model, and initialize the training information;

[0075] (2) Load the training image files, perform preprocessing and convert them into array form;

[0076] (3) Load the image list img and the label one-hot code label;

[0077] (4) Rearrange label in the order of image loading to obtain Truth;

[0078] (5) Divide the training set and the test set and shuffle them randomly;

[0079] In step (1), the initialization training includes the number of iterations, batch size, learning rate, image size, etc.

[0080] In step (2), the preprocessing includes: data augmentation, translation, rotation, resizing, and normalization.

[0081] In step (3), the image list and the label one-hot code are obtained from step 3 of the second aspect.

[0082] The reason for step (4) is as follows: ① The training set image list img_train = [1.jpg, 10.jpg, 100.jpg,...], while img obtained from step 3 of the second aspect is [1.jpg, 2.jpg, 3.jpg...]. The order of image loading during training may be inconsistent with the original order in the img list; ② It is possible to use a method of dividing data and select some images for training instead of all in the img list. Directly using img obtained from step 3 of the second aspect may result in image-label mismatch.

[0083] The specific method of step (4) is as follows: ① Traverse all the images in the training set folder to obtain the training set image list img_train; ② For each image in img_train, find the position index index of the image name in img_train; ③ Use the index index to search for its corresponding label in the label list obtained in step 3 of the second aspect, and then reconstruct the label list truth in the order of the training set image list img_train.

[0084] In step (5), the method of dividing the data set is as follows: randomly select 80% of the data set for training, and the remaining 20% for testing. The purpose of shuffling is to reshuffle the data so that the arrangement of the data has a certain degree of randomness, and the possibility of any type of data is the same when reading in order.

[0085] 2. Model configuration:

[0086] (1) Select the backbone model and configure the output;

[0087] (2) Configure the optimizer and loss function;

[0088] In step (1), the backbone model can be Vgg, Resnet, Inception, etc. The end of the model is a fully connected layer, and the number of output nodes of the model is equal to the number of attribute categories, which is 27. The activation function at the end of the model selects sigmoid, and its function is to non-linearly transform each output value of the network, converting a scalar number to the range [0, 1], which is convenient for binary classification to determine whether the attribute exists.

[0089] In step (2), the Adam optimization method is selected for the optimizer, and the binary cross-entropy loss is selected for the loss function. The reason is that multi-label classification treats each predicted label as an independent Bernoulli distribution to evaluate each output node separately.

[0090] 3. Model training and saving:

[0091] The input for training the model is img and Truth obtained in step (4) of the data preparation unit - 1 of the model construction and training unit in the third aspect.

[0092] (2) Model testing and evaluation unit:

[0093] 1. Model testing:

[0094] (1) Load the model and label file;

[0095] (2) Process the image to be tested and input it into the model;

[0096] (3) Mark the attribute index according to the model output and the set threshold;

[0097] (4) Obtain the prediction result img_pre according to the attribute index;

[0098] The model in the step (1) is obtained by the third aspect (1) model construction and training unit - 3 model training and saving, and the label file is obtained by the step 3 of the second aspect

[0099] In the step (2), the image processing operations include: resizing, normalizing, and converting to an array form.

[0100] In the step (3), the image in array form processed by the step (2) is input into the model loaded in the step 1. The output of the model is between [0, 1], which is the predicted value of the non-linear processing of the sigmoid activation function at the end of the network. Set a certain threshold (usually set to 0.5), and define that if the output is greater than the threshold, the task attribute exists, otherwise it does not exist. Denote the index corresponding to the value greater than the threshold in the output list as idx.

[0101] In the step (4), according to the index value idx obtained in the step (3), obtain the prediction result pre. Specifically, define a list of length 27 all of 0, and at the index idx, set the value of this list to 1, that is, obtain the prediction result of each attribute.

[0102] 2. Image-level evaluation:

[0103] (1) Load the attribute prediction value img_pre of the image;

[0104] (2) Load the true value img_truth of the label of the image;

[0105] (3) Calculate the evaluation index of each image;

[0106] In the step (1), load the image prediction value img_pre obtained in the third aspect (2) model testing and evaluation unit - 1 model testing - step (4), which is a list of length equal to 27, and the list elements are 0 or 1.

[0107] In the step (2), load the Label information obtained by the step 3 of the second aspect, that is, the true value of the image label, which is a list of length equal to 27, and the list elements are 0 or 1.

[0108] In step (3), the functions provided in sklearn.metrics of Python machine learning, namely accuracy_score, precision_score, and recall_score, can be used to calculate the metrics. The inputs of the functions are the predicted values img_pre obtained from step (1) and the true values img_truth loaded in step (2).

[0109] 3. Evaluation at the attribute level:

[0110] (1) Concatenate the image predicted values to obtain the attribute predicted value n_pre;

[0111] (2) Concatenate the image true values to obtain the attribute true value n_truth;

[0112] (3) Construct the predicted values att_pre and true values att_truth for each attribute;

[0113] (4) Calculate the evaluation metrics for each attribute;

[0114] In step (1), concatenate the predicted values img_pre of the n images obtained in step (1) of the image-level evaluation in the third aspect (2) model testing and evaluation unit - 2 to obtain the attribute predicted value n_pre, which is a list with a length equal to n, and the list elements are predicted values img_pre with a length of 27.

[0115] In step (2), concatenate the true label values img_truth of the n images obtained in step (2) of the image-level evaluation in the third aspect (2) model testing and evaluation unit - 2 to obtain the attribute true value n_truth, which is a list with a length equal to n, and the list elements are true label values img_truth with a length of 27.

[0116] In step (3), regard n_pre and n_truth obtained in step (1) and step (2) as an n-row and m-column matrix, where n is the number of images and m is the number of attributes, and m is equal to 27. Then each column of n_pre represents the predicted value att_pre of the corresponding attribute in this column, att_pre is a list with a length equal to n, and the list elements are 0 or 1; each column of n_truth represents the true value att_truth of the corresponding attribute in this column, att_truth is a list with a length equal to n, and the list elements are 0 or 1.

[0117] In step (4), functions provided in sklearn.metrics in Python machine learning, namely accuracy_score, precision_score, and recall_score, can be used to calculate the metrics. The inputs of the functions are the predicted values att_pre and the true values att_truth of each attribute obtained in step (3).

[0118] Fourthly, as Figure 6 , the present invention provides an elevator passenger attribute recognition device, including an elevator image data acquisition module, an elevator passenger attribute data set establishment and processing module, an elevator passenger attribute recognition and evaluation module, and a result display and saving module that are communicatively connected in sequence;

[0119] The elevator image data acquisition module is used to obtain the in-car monitoring video and intercept the passenger images by using a deep learning algorithm.

[0120] The elevator passenger attribute data set establishment and processing module is used to establish an elevator image folder and a label csv file, traverse and proofread them to make them correspond one by one, extract the image and label lists respectively, and convert the labels into binary one-hot codes.

[0121] The elevator passenger attribute recognition module is used to configure, train, test the deep neural network model and quantitatively evaluate the prediction results.

[0122] The result display and saving module is used to construct a prediction result text dictionary, load various evaluation metrics, print and save the prediction evaluation results. The display module includes the following steps:

[0123] 1. Load the image and construct the background;

[0124] 2. Construct a prediction result text dictionary;

[0125] 3. Load various evaluation metrics;

[0126] 4. Print and save the prediction evaluation results;

[0127] In step 1, load the image tested in the model test of the third aspect (3), adjust the size and generate a black background of the same size.

[0128] In step 2, organize the attribute prediction results in the form of a dictionary, that is, key-value pairs. The key is 27 attribute categories, which are loaded from the pkl file obtained in step 5 of the second aspect, and the value is the output of the image sent into the model, representing the probability of each attribute in all possible target categories, which is loaded from the model test of the third aspect (3).

[0129] In step 3, load the evaluation index results calculated in the evaluation step of the third aspect (4) at the image level.

[0130] In step 4, print the result text constructed in step 2 and the evaluation indexes loaded in step 3 on the black background established in step 1 in sequence, then perform a horizontal splicing with the original image to be measured loaded in step 1, and finally save it in a specified path.

[0131] Fifth aspect, the present invention provides an intelligent multi-label attribute classification system for elevator car passengers, including a monitoring camera, a video server, and a personal computer that are communicatively connected in sequence. Among them, the monitoring camera is arranged at the top inside the car and the monitoring field of view covers the area where the car door is located;

[0132] The monitoring camera is used to transmit the acquired video stream to the video server;

[0133] The video server is used to perform digital conversion on the video stream to obtain the in-car monitoring video;

[0134] The personal computer is used to obtain the in-car monitoring video from the video server and execute the intelligent multi-label attribute classification method for elevator passengers as described in the third aspect to obtain the recognition result of the passenger appearance characteristics.

[0135] Sixth aspect, as Figure 7 , the present invention provides a computer device, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the label data processing method as described in the second aspect or the intelligent multi-label attribute classification method for elevator passengers as described in the third aspect.

[0136] Seventh aspect, the present invention provides a computer-readable storage medium, on which instructions are stored. When the instructions are run on a computer, the label data processing method as described in the second aspect or the intelligent multi-label attribute classification method for elevator passengers as described in the third aspect is executed.

[0137] Eighth aspect, the present invention provides a computer program product containing instructions. When the instructions are run on a computer, the computer is made to execute the label data processing method as described in the second aspect or the intelligent multi-label attribute classification method for elevator passengers as described in the third aspect.

[0138] Embodiment 2:

[0139] Embodiment 2 is a preferred example of Embodiment 1.

[0140] As Figure 1As shown, the intelligent multi-label attribute classification method for elevator passengers based on deep learning includes:

[0141] S1. Obtain the elevator car monitoring video and transcode it for smooth playback.

[0142] The elevator monitoring video is a monitoring video collected for the target elevator car, which can be a real-time collected monitoring video or a historically collected monitoring video stored in, for example, a video server or a local memory. Therefore, the specific method of obtaining it can include, but is not limited to: receiving the real-time monitoring video from a monitoring camera, or reading the historical monitoring video from the video server or the local memory.

[0143] S2. Invoke the target detection algorithm to detect people and intercept the passenger images.

[0144] As Figure 2 shown, the people detection unit invokes the deep learning method to detect people in the video stream. Optional target detection algorithms include: single-stage detectors such as YOLO and SSD that directly perform classification and regression based on feature extraction, and two-stage detectors such as Fast R-CNN based on candidate regions. The image interception unit is used to intercept the people detection regression box in the elevator video stream and save it to the dataset path.

[0145] S3. Establish an elevator passenger attribute dataset and traverse and proofread the images and labels.

[0146] As Figure 3 shown, first obtain the elevator image dataset. It includes an image folder and a label csv file, where the image names are named with increasing Arabic numerals in the order of interception. The first column of the csv file is the image name, and the 2nd - 12th columns are the corresponding label information. Secondly, process the dataset, traverse and proofread the images and labels to make them correspond one by one. It includes the following steps: read the image and label lists, obtain the missing files by comparing the intersections, delete the incomplete difference set files, and delete the duplicate label files.

[0147] In step S3, the label attributes include the following categories: gender (male, female); age (child, teenager, adult, elderly); headgear (without hat, with hat, with helmet); hair length (unidentifiable, bald, short hair, long hair); whether white hair (non-white hair, white hair); whether glasses (without glasses, with glasses); occupation (courier, cleaner, resident); whether holding a child (not holding a child, holding a child); whether playing with mobile phone (not playing with mobile phone, playing with mobile phone); whether carrying a backpack (not carrying a backpack, carrying a backpack); whether carrying a handbag (carrying a handbag, carrying a handbag), a total of 11 types.

[0148] S4. Extract the image and label lists respectively and convert the labels into one-hot codes.

[0149] As Figure 3 shown, first, extract the image and label list respectively. That is, read the first column of the csv file as the image list, the length of which is equal to the number of images, and the list content is the image name; read the row corresponding to each image index as the label list, the length of which is equal to the number of attribute types, i.e., 11, and the list content is the text label. Secondly, convert the labels into binary one-hot codes. Binarize the label list obtained in the third step. Specifically, let the attribute be recorded as 1 and the non-existent be recorded as 0, convert the label list into one-hot codes, and save the binarized multi-labels as a pkl file for easy loading and printing of the attribute recognition results. After being binarized into one-hot codes, expand the attributes of each category, namely "not wearing a hat, not wearing glasses, not holding a child, not carrying a handbag, not playing with a mobile phone, not carrying a backpack, cleaning staff, bald / with a bald head, hair unrecognizable, female, child, resident, courier, adult, wearing a helmet, wearing a hat, wearing glasses, carrying a handbag, playing with a mobile phone, male, white hair, short hair, old person, carrying a backpack, long hair, teenager, non-white hair", a total of 27 kinds, so the length of the one-hot code label is 27.

[0150] S5. Load the images and preprocess them, load the attribute labels and rearrange them.

[0151] As Figure 4 shown, the step S5 corresponds to the data preparation unit. First step: Configure the dataset and the path of the pre-trained model, and initialize the training information, including the number of iterations, batch size, learning rate, image size, etc. Second step: Load the training image files, preprocess them and convert them into array form. The preprocessing includes: data augmentation, random translation, rotation, resizing, normalization. Third step: Load the image list and label one-hot codes obtained in the step S4. Fourth step: Rearrange the labels according to the order of image loading to obtain the true values. The specific steps include: traverse all the images in the training set folder to obtain the training set image list; for each image, find the position index of the image name in the training set image list; use the index to find its corresponding label in the label list obtained in the step S4, and then re-sort the label list according to the loading order of the training set image list to ensure one-to-one correspondence with the images.

[0152] S6. Divide the dataset, configure the parameters of the deep neural network model and train.

[0153] In the step S6, randomly select 80% of the dataset for training, and the remaining 20% for testing. Shuffle the data to make the arrangement of the data have a certain degree of randomness, so that the possibility of any type of data is the same when reading in order.

[0154] As Figure 4As shown in the figure, the model configuration unit includes configuring the backbone model parameters, the optimization function, and the loss function. The backbone model can be selected from Vgg, Resnet, Inception, etc. The end of the model is a fully connected layer, and the number of output nodes of the model is equal to the number of attribute categories, which is 27. The activation function at the end of the model is selected as sigmoid, and its function is to non-linearly process each output value of the network, converting a scalar number to the range [0, 1], which is convenient for binary classification to determine whether the attribute exists. Since multi-label classification treats each predicted label as an independent Bernoulli distribution to evaluate each output node separately, the Adam optimization method is selected as the optimizer, and the binary cross-entropy loss is selected as the loss function. The input of the trained model is the images loaded in sequence in step S5 and the attribute labels reordered in step S5.

[0155] S7. Model prediction, calculating evaluation metrics at the image level and the attribute level.

[0156] As Figure 5 shown in the figure, the specific implementation steps of model testing are as follows: Load the model trained in step S6 and the test images obtained in step S5. After preprocessing such as resizing, normalizing, and converting to array form, the images are input into the model. The output of the model is between [0, 1], which is the predicted value non-linearly processed by the sigmoid activation function at the end of the network. Set a certain threshold (usually 0.5). If the output is greater than the threshold, it is considered that the attribute exists; otherwise, it does not exist. Record the indices corresponding to the values greater than the threshold in the output list. Define a list of length 27 that is all 0. At the indices, set the values of the list to 1, and thus obtain the prediction results for each attribute. The attribute prediction result for each image is a list of length 27, and the list elements are 0 or 1.

[0157] As Figure 5 shown in the figure, the specific implementation steps of image-level evaluation are as follows: Load the image prediction list obtained in step S7 and the true label list obtained in step S5 to calculate three evaluation metrics for each image. The calculation methods that can be selected include the functions provided in sklearn.metrics in Python machine learning: accuracy_score, precision_score, recall_score, etc.

[0158] As Figure 5As shown, the specific implementation steps for the property-level evaluation are as follows: For n images, the image prediction list obtained in step S7 is concatenated to obtain the property prediction value, and the true label list obtained in step S5 is concatenated to obtain the property true value. Both the property prediction value and the property true value are matrices with n rows and m columns, where n is the number of images and m is the number of properties, and m is equal to 27. Each column of the matrix represents the prediction value or true value of a property, and each row represents the prediction value or true value of an image. Each column of the property prediction value matrix is extracted to construct the prediction value of each property, and each column of the property true value matrix is extracted to construct the true value of each property, which serves as the input for calculating the evaluation metrics and is used to calculate three evaluation metrics for each property. The available calculation methods include functions provided in sklearn.metrics in Python machine learning: accuracy_score, precision_score, recall_score, etc.

[0159] S8. Construct a background and prediction dictionary, print the property recognition and evaluation results, and save them.

[0160] In step S8, the display module includes the following steps: Load the images tested in step 5, adjust the size, and generate a black background of the same size. Organize the property prediction results obtained in step 7 in the form of a dictionary, i.e., key-value pairs, where the key is one of the 27 property categories, loaded from the pkl file obtained in step S4, and the value is the output of the image fed into the model, representing the probability of each property in all possible target categories, loaded from the model test in step S7. Load the evaluation metric results calculated for the image-level evaluation obtained in step S7. Print the constructed result text and the loaded evaluation metrics on the established black background in sequence, then horizontally concatenate them with the original image to be tested for easy comparison, and finally display and save them at the specified path.

[0161] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be considered as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.

[0162] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A passenger multi-label classification system based on deep learning, characterized in that, it includes: Image data acquisition module: Obtain the surveillance video, transcode and play it, and call the object detection algorithm to detect people and intercept the passenger images; Attribute data set establishment and processing module: Establish a passenger attribute data set, traverse and proofread the images and labels, extract the image and label lists respectively, and convert the labels into one-hot codes; Passenger attribute recognition module: First load the images and preprocess them, load the attribute labels and rearrange them, then divide the data set, configure the deep neural network model parameters and train, and then perform model prediction, and calculate evaluation indicators at the image level and the attribute level; Result display and saving module: Construct the background and prediction dictionaries, print the attribute recognition and evaluation results and save them; The image data acquisition module includes a surveillance video acquisition unit, a person detection unit, and an image interception unit that are sequentially communicatively connected; The surveillance video acquisition unit uses the camera installed on the top of the elevator to collect the surveillance video in real time, transcode it and play it smoothly; The person detection unit calls the deep learning method to detect the people in the video stream; The image interception unit is used to intercept the person detection regression box in the video stream and save it to the data set path; The attribute data set establishment and processing module includes: Establish an elevator image data set, including an image folder and a label csv file; Use the passenger attribute data set establishment and processing method to traverse and proofread the images and labels to make them correspond one by one; Extract the image and label lists respectively, and convert the text labels into binary one-hot codes; The passenger attribute recognition module includes: Model construction and training unit, used to prepare training data, configure network parameters and train the deep learning model; Model testing and evaluation unit, used to evaluate the model performance at the image level and the attribute level respectively, and the evaluation indicators include accuracy, precision and recall rate.

2. A passenger multi-label classification method based on deep learning, characterized in that, it includes: Image data acquisition step: Obtain the surveillance video, transcode and play it, and call the object detection algorithm to detect people and intercept the passenger images; Attribute data set establishment and processing step: Establish a passenger attribute data set, traverse and proofread the images and labels, extract the image and label lists respectively, and convert the labels into one-hot codes; Passenger attribute recognition step: First load the images and preprocess them, load the attribute labels and rearrange them, then divide the data set, configure the deep neural network model parameters and train, and then perform model prediction, and calculate evaluation indicators at the image level and the attribute level; Result display and saving step: Construct the background and prediction dictionaries, print the attribute recognition and evaluation results and save them; The image data acquisition step includes: Surveillance video acquisition step: Use the camera installed on the top of the elevator to collect the surveillance video in real time, transcode it and play it smoothly; Person detection step: Call the deep learning method to detect the people in the video stream; Image interception step: Intercept the person detection regression box in the video stream and save it to the data set path; The attribute data set establishment and processing step includes: Establish an elevator image data set, including an image folder and a label csv file; Method for establishing and processing using a passenger attribute dataset, traversing and proofreading images and labels to make them correspond one by one; Extract the image and label lists respectively, and convert the text labels into binary one-hot codes; The passenger attribute recognition steps include: Model construction and training steps: Prepare training data, configure network parameters, and train a deep learning model; Model testing and evaluation steps: Evaluate the model performance at the image level and the attribute level respectively, and the evaluation metrics include accuracy, precision, and recall.

3. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method described in claim 2.

4. A deep learning-based passenger multi-label classification device, characterized in that, comprising: a controller; The controller includes the computer-readable storage medium storing the computer program described in claim 3, and when the computer program is executed by a processor, it implements the steps of the deep learning-based passenger multi-label classification method described in claim 2; or, the controller includes the deep learning-based passenger multi-label classification system described in claim 1.

Citation Information

Patent Citations

  • Escalator passenger falling detection algorithm based on deep learning

    CN108805093A

  • Deep learning-based escalator passenger fall detection method

    CN108805093B

  • Escalator passenger retrograde motion detection method based on deep learning

    CN111144247A

  • A Deep Learning-Based Method for Detecting Passengers Going Against Traffic on Escalators

    CN111144247B

  • Escalator retention detection method and system based on Gaussian mixture model

    CN111259718A