Elevator sensorless control system and method
By using binocular vision units and deep learning algorithms for face detection and feature extraction in the intelligent elevator system, the problems of long time and high false recognition rate in the existing elevator system are solved, fast and accurate identity authentication is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202311586817.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-11-27
AI Technical Summary
The existing smart elevator system takes a long time during the face recognition process, with a high misrecognition rate, which affects the user experience.
The binocular vision unit is combined with deep learning algorithms to obtain video streams through the binocular camera, face detection and feature extraction are performed, and face features are compared using databases to achieve fast and accurate identity authentication.
It significantly improves the face recognition rate and accuracy of the elevator, improves the user experience, and reduces the misrecognition rate.
Smart Images

Figure CN117623031B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of elevator control, and particularly relates to an elevator touchless control system and method. Background Art
[0002] At present, society is moving towards an era of rapid development of informatization, networking, and digitalization. Information security has become particularly important and has attracted more and more attention. Information security involves all aspects of society. Therefore, how to effectively protect information security and how to authenticate identities conveniently, quickly, and effectively has become a hot topic in society. With the advent of the 5G era with the rise of artificial intelligence technology, many problems that were difficult to solve before have received more technical support and solutions. However, with the rapid development of face recognition technology, its application in actual production and life has become more and more popular and mature, bringing great convenience and better experience to people's lives. Intelligent elevators transmit pictures of almonds in real time through cameras, detect the faces of pedestrians from the pictures to verify identities, and display the issued elevator control instructions and simulated elevator states on the user interface. Usually, the measurement of distance and the processes of face detection and face recognition take a long time, and there will also be problems of misidentification. Summary of the Invention
[0003] In view of this, the present invention provides an elevator touchless control system and method that can improve the face recognition rate and accuracy to solve the above-mentioned existing technical problems, and specifically adopts the following technical solutions to achieve.
[0004] In a first aspect, the present invention provides an elevator touchless control system, including:
[0005] A binocular vision unit, configured to obtain a video stream from a binocular camera, obtain image information from the video stream, input the image information into a face detection unit and a face recognition unit for processing to obtain a face image, and perform binocular ranging on the face in the video according to the face image;
[0006] A face detection unit, configured to detect a face from the face image to obtain the position and size of the face, wherein a deep learning detection algorithm is used for face detection;
[0007] A face recognition unit, configured to extract features of the face in the detected face image to obtain face features. Among them, a feature vector corresponding to the face is extracted through convolution, activation, and pooling operations, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed a preset threshold, there is no such person in the database. If it exceeds the preset threshold, the face identity information corresponding to the maximum value is taken;
[0008] A display unit for displaying the results of face recognition. The recognized face is marked with a square, and the identity information and the status of the elevator are given. Among them, the elevator status includes the opening and closing of the elevator door, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state;
[0009] A database unit for storing face feature vectors, the identity information corresponding to the face feature vectors, and the status of the elevator.
[0010] As an optimization of the above technical solution, the display unit includes an image preprocessing module for processing the elevator door image. The image preprocessing module uses the Gaussian filtering algorithm to perform image smoothing on the elevator door image to obtain a smoothed image g(x,y), and the corresponding expression is g 1 (x,y) = g(x,y) * G(x,y), where σ represents the standard variance of Gaussian filtering;
[0011] Taking a 3×3 template as an example, within the template range, calculate the gray difference grd(x,y) between the current pixel point g(x,y) and other pixels, and the corresponding expression is
[0012] Let the Gaussian smoothing factor α = 1 - grd(x,y) / 255, and smooth the image through the Gaussian smoothing factor to obtain the improved smoothed image g 2 (x,y), and the expression is g 2 (x,y) = αg 1 (x,y) + (1 - α)g(x,y). When the image pixel is in the edge area, the greater the gray difference grd(x,y) between the pixels, the smaller the Gaussian smoothing factor α.
[0013] As an optimization of the above technical solution, the display unit further includes an edge detection module for processing the elevator door image. The edge detection module uses an adaptive Canny edge detection algorithm, and calculates the gradient magnitude and gradient direction of the smoothed image g 2 (x,y) through solving the finite difference mean value. The expressions are where P x [x,y] represents the partial derivative function in the x direction, and P y [x,y] represents the partial derivative function in the y direction. The gradient magnitude is The gradient direction is θ[x,y] = arctan(P y [x,y] / P x [x,y]), add the gradient magnitudes in the 45° direction and the 135° direction, and the expression of the obtained total gradient magnitude is where P x [x,y], Py [x, y], P xy [x, y] and P yx [x, y] respectively represent the partial derivative functions in the x-axis direction, y-axis direction, 45° direction, and 135° direction;
[0014] The Canny operator realizes edge detection by setting high and low thresholds. If the gradient magnitude of a certain pixel is higher than the high threshold, then this pixel is an edge point; if the gradient magnitude of a certain pixel is lower than the low threshold, then this pixel is not an edge point.
[0015] As an optimization of the above technical solution, the display unit further includes an edge screening module. The edge screening module extracts detection identification point information using image representation. The detection identification points include coded identification points and non-coded identification points. The coded identification points and non-coded identification points are circles of the same size and appear elliptical after CCD imaging. The execution process of the edge screening module includes:
[0016] The elliptical contour is a convex closed contour. Compare the distance between the starting pixel point and the ending pixel point of the curve segment after edge tracking. If the distance is less than the threshold, it is determined that the contour is closed; otherwise, it is noise and is removed;
[0017] Select a projection angle, that is, the angle between the normal direction of the plane where the identification point is located and the projection direction, within 0° - 60° to take pictures of the elevator door. Within this angle range, the contour perimeter L and elliptical area S of the identification points in the elevator door image respectively satisfy where L min 、L max respectively represent the minimum and maximum ranges that the contour perimeter of the identification points should satisfy within this angle range, and S min 、S max respectively represent the minimum and maximum ranges that the contour area of the identification points should satisfy within this angle range;
[0018] Using the geometric characteristics of the ellipse, the target edge contour is screened. The expression for screening the ellipse contours that meet the requirements within the range of 0° - 60° through the aspect ratio of the ellipse's major and minor axes, shape factor R, and circularity C of the ellipse is where a and b respectively represent the major and minor axes of the ellipse, and are required to be less than the threshold M.
[0019] As an optimization of the above technical solution, the contrast between the foreground color and the background color of the detection identification points is strong. According to the gray-scale characteristics of the marked points, the contours after edge detection are screened. The central area of the preset detection identification points is the foreground color E 1 , the gray-scale value of the pixel points is close to white, and M f represents the average value of the pixel points within the range of E 1 ; the annular area is the background color E 2 , the gray-scale value of the pixel points is close to black, and Mb Represents E 2 The average value of pixels within the range, then M f and M b should satisfy where M f represents the threshold for distinguishing the foreground color and background color of the image, and ΔM t represents the minimum threshold that the difference between the foreground color and the background color should satisfy.
[0020] As an optimization of the above technical solution, the face recognition unit includes a model selection module. The model selection module uses recognition signals and verification signals, and the total loss is the weighted sum of the losses of the recognition signal and the verification signal. The weights are dynamically adjusted during the training process. The expressions of their respective loss functions are where f represents the feature vector, θ represents the network parameters, t represents the target category, p represents the target probability distribution, represents the predicted probability distribution. When i = 1, p i = 1. When the two samples are of the same person, y is 1, otherwise it is -1;
[0021] The expression for training using an optimizer is where m t and v t represent the first and second order momenta, and β 1 and β 2 represent the momentum parameters, represents the correction value, and W t represents the parameters of the network at time t.
[0022] As an optimization of the above technical solution, the face recognition unit further includes a model evaluation module. The model evaluation module uses accuracy, precision, and recall as evaluation indicators for the generalization ability of the model. The authenticity of the model prediction result is determined by a confidence threshold. The confidence threshold is the IoU threshold. The IoU threshold is set to 0.5. If the confidence is greater than 0.5, it is true; if the confidence is less than 0.5, it is false. Accuracy: Precision: Recall:
[0023] As an optimization of the above technical solution, the execution process of the face detection unit includes:
[0024] Detect a face and extract the feature vector of this face;
[0025] Compare this feature vector with all the feature information in the label, find the feature vector with the closest Euclidean distance to it, use the distance between these two vectors as the key value, and use the identity information corresponding to the corresponding feature vector in the label as the value, and save them in pairs to the temporary dictionary of the database unit;
[0026] According to the detected number of people, preset as n, extract the values of the first n values with the smallest key values in the temporary dictionary as the recognized information, and send out the command signal for setting the floor to complete the elevator touchless control.
[0027] As an optimization of the above technical solution, the binocular vision unit includes a calibration module. The calibration module obtains the position information between the binocular cameras through binocular calibration. The positional relationship can be described by the translation matrix P and the rotation matrix T. By solving the translation and rotation matrices between the left and right camera coordinates and the world coordinate, the expression for the positional relationship between the left and right camera coordinates can be obtained as where P l represents the coordinates of P in the left camera coordinate system in the world coordinate system. Through the rotation matrix P l and the translation matrix T l the conversion between the left camera coordinate and the world coordinate is realized; P r represents the coordinates of P in the right camera coordinate system in the world coordinate system. Through the rotation matrix R r and the translation matrix T r the conversion between the right camera coordinate and the world coordinate is realized.
[0028] In the second aspect, the present invention also provides an elevator touchless control method, including the following steps:
[0029] Obtain the video stream from the binocular camera, obtain the image information from the video stream and input the image information into the face detection unit and the face recognition unit for processing to obtain the face image, and perform binocular ranging on the faces in the video according to the face image;
[0030] Detect the face from the face image to obtain the position and size of the face. Among them, a deep learning detection algorithm is used for face detection;
[0031] Extract the face features from the face in the detected face image. Among them, the feature vector corresponding to the face is extracted through convolution, activation and pooling operations, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed the preset threshold, there is no such person in the database. If it exceeds the preset threshold, the face identity information corresponding to the maximum value is taken;
[0032] Display the face recognition result. The recognized face is marked with a box, and the identity information and the status of the elevator are given. Among them, the elevator status includes the opening and closing of the elevator door, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state;
[0033] Store the face feature vector, the identity information corresponding to the face feature vector and the status of the elevator.
[0034] The present invention provides an elevator passive sensing control system and method. By obtaining a video stream from a binocular camera, image information is obtained from the video stream and input into a face detection unit and a face recognition unit for processing to obtain a face image. Binocular ranging is performed on the face in the video according to the face image. The face is detected from the face image to obtain the position and size of the face. Feature extraction is performed on the face in the detected face image to obtain face features. The face recognition result is displayed, and the recognized face is marked with a square box, and identity information and the state of the elevator are given. The face feature vector, the identity information corresponding to the face feature vector, and the state of the elevator are stored, so that ranging can be performed on the detected face and the face can be quickly recognized, thereby improving the face recognition rate and accuracy of the elevator and also improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a structural block diagram of the elevator passive sensing control system of the present invention;
[0037] Figure 2 It is a flowchart of the elevator passive sensing control method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0039] Refer to Figure 1 , the present invention provides an elevator passive sensing control system, including:
[0040] A binocular vision unit, configured to obtain a video stream from a binocular camera, obtain image information from the video stream, and input the image information into a face detection unit and a face recognition unit for processing to obtain a face image, and perform binocular ranging on the face in the video according to the face image;
[0041] A face detection unit, configured to detect a face from the face image to obtain the position and size of the face, wherein a deep learning detection algorithm is used for face detection;
[0042] The face recognition unit is used to extract the features of the face in the detected face image to obtain the face features. Among them, through convolution, activation, and pooling operations, the feature vector corresponding to the face is extracted, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed the preset threshold, there is no such person in the database. If it exceeds the preset threshold, the face identity information corresponding to the maximum value is taken;
[0043] The display unit is used to display the face recognition result. The recognized face is marked with a square frame, and the identity information and the state of the elevator are given. Among them, the elevator state includes the opening and closing of the elevator door, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state;
[0044] The database unit is used to store the face feature vectors, the identity information corresponding to the face feature vectors, and the state of the elevator.
[0045] In this embodiment, the display unit includes an image preprocessing module for processing the elevator door image. The image preprocessing module uses the Gaussian filtering algorithm to perform image smoothing on the elevator door image to obtain the smoothed image g(x, y), and the corresponding expression is g 1 (x, y) = g(x, y) * G(x, y), where σ represents the standard variance of Gaussian filtering; taking a 3×3 template as an example, within the template range, the gray difference grd(x, y) between the current pixel point g(x, y) and other pixels is calculated, and the corresponding expression is Let the Gaussian smoothing factor α = 1 - grd(x, y) / 255, and the image is smoothed through the Gaussian smoothing factor to obtain the improved smoothed image g 2 (x, y), and the expression is g 2 (x, y) = αg 1 (x, y) + (1 - α)g(x, y). When the image pixel is in the edge area, the gray difference grd(x, y) between pixels is also larger, and the Gaussian smoothing factor α is smaller. The binocular vision unit includes a calibration module. The calibration module obtains the position information between the binocular cameras through binocular calibration. The position relationship can be described by the translation matrix P and the rotation matrix T. By solving the translation and rotation matrices between the left and right camera coordinates and the world coordinate, the expression of the position relationship between the left and right camera coordinates can be obtained as where P l represents the coordinates of P in the left camera coordinate system in the world coordinate system. Through the rotation matrix R l and the translation matrix T l realize the conversion between the left camera coordinate and the world coordinate; P rDenote the coordinates of P in the right camera coordinate system in the world coordinate system, through the rotation matrix R r and the translation matrix T r to achieve the conversion between the right camera coordinate and the world coordinate.
[0046] It should be noted that the binocular vision unit acquires video streams, and inputs the pictures into the face detection unit, face recognition unit and display unit every fixed frame. After preprocessing the pictures, the face detection unit sends them into the detection network to detect faces. After detecting faces, it provides face position information for the face recognition unit and display unit. At the same time, a specific point is found on the corresponding face frames of the left and right cameras, and its pixel coordinates on the left and right cameras are transmitted to the binocular vision unit for distance measurement. According to the face position information, the face recognition unit extracts the face from the picture, aligns the face pictures and then sends them into the feature extraction network for feature extraction, and compares the extracted face feature vectors with the face feature vectors in the database unit to identify user information. After identifying the information, it is sent to the display unit for display, and the control instruction for the elevator is obtained based on the identified information of the elevator-riding user and the current state of the elevator, that is, it is displayed on the display unit. For the users whose information cannot be identified and there is no such user information in the database, the extracted face features, the manual operation information of the corresponding users and the set floor information are stored in the database.
[0047] It should be understood that when the elevator is running, the system receives real-time video input, performs face detection on each frame of the image of the left camera. When a face is detected, the face position information is sent to the recognition network for subsequent processing. At the same time, face detection is performed on the image of the right camera, and the pixel coordinates of a certain feature point, such as the center point of the face, corresponding to the face frames in the left and right image frames are obtained. According to the binocular ranging principle, the distance of this person from the elevator door is calculated. The binocular cameras are installed on one side of the elevator door at a moderate height. When it is too low, it is easy for the people behind to be blocked by the people in front. When it is too high, the frontal information of the detected face will be reduced. When ranging, it can be considered that the Z coordinate value of a point in space is the distance of this point from the elevator door. Receive the face distance information and identity information, and at the same time read the real-time state of the elevator from the traditional elevator control system, and make a control instruction response for the elevator by integrating the three pieces of information. When the user is no more than two meters away from the elevator door, if the elevator is not on the first floor, an up call instruction for the first floor is issued. If the elevator is on the first floor, an instruction to open the elevator door is issued. After the elevator door is opened, the floor is set according to the identified user information, and then the control instruction is sent to the control system. The control system comprehensively processes information such as the inside and outside call signals, sensor signals and the instruction signals of this system of the elevator, and performs non-sensing control on the movement of the elevator, so as to range the detected face and quickly identify the face, improving the face recognition rate and accuracy of the elevator, and also improving the user experience.
[0048] Optionally, the display unit further includes an edge detection module for processing the elevator door image. The edge detection module uses an adaptive Canny edge detection algorithm to calculate the smoothed image g by solving the finite difference mean 2 The expressions for the gradient magnitude and gradient direction of g(x,y) are where P x [x,y] represents the partial derivative function in the x direction, and P y [x,y] represents the partial derivative function in the y direction. The gradient magnitude is The gradient direction is θ[x,y] = arctan(P y [x,y] / P x [x,y]). By increasing the gradient magnitudes in the 45° and 135° directions, the expression for the total gradient magnitude obtained is where P x [x,y], P y [x,y], P xy [x,y], and P yx [x,y] represent the partial derivative functions in the x-axis direction, y-axis direction, 45° direction, and 135° direction respectively;
[0049] The Canny operator realizes edge detection by setting high and low thresholds. If the gradient magnitude of a certain pixel point is higher than the high threshold, then this pixel point is an edge point; if the gradient magnitude of a certain pixel point is lower than the low threshold, then this pixel point is not an edge point.
[0050] In this embodiment, the display unit further includes an edge screening module. The edge screening module extracts detection identification point information using image representation. The detection identification points include coded identification points and non-coded identification points. The coded identification points and non-coded identification points are circles of the same size and appear elliptical after CCD imaging. The execution process of the edge screening module includes: The elliptical contour is a convex closed contour. Compare the distance between the starting pixel point and the ending pixel point of the curve segment after edge tracking. If the distance is less than the threshold, it is determined that the contour is closed, otherwise it is noise and is eliminated; Select a projection angle, that is, the angle between the normal direction of the plane where the identification point is located and the projection direction, within 0° - 60° to photograph the elevator door. Within this angle range, the contour perimeter L and elliptical area S of the identification points in the elevator door image respectively satisfy where L min , L max respectively represent the minimum and maximum ranges that the contour perimeter of the identification points should satisfy within this angle range, and S min , S maxrespectively represent the minimum and maximum ranges that the contour area of the identification point should satisfy within this angular range; the geometric characteristics of the ellipse will be used to screen the target edge contour, and the ellipse contours that meet the requirements within the range of 0°-60° will be screened. The expressions for screening the contour by the aspect ratio of the major and minor axes of the ellipse, the shape factor R, and the roundness C of the ellipse are where a and b respectively represent the major and minor axes of the ellipse, and are required to be less than the threshold M.
[0051] It should be noted that edge detection refers to detecting edge points with gray-scale mutations or texture changes from an image, which is an important prerequisite for subsequent image feature recognition and extraction. The quality of edge point extraction will directly affect the accuracy of identification extraction. The Canny operator uses high and low thresholds to detect strong and weak edges respectively, is not easily affected by noise, and can detect real weak edges. By the OSTU algorithm to condition the high and low thresholds of Canny edge detection, the size of the preset elevator door gray-scale image is M*N, f(x,y) represents the image pixel at a certain point in the image, h(x,y) represents the average gray-scale value within the neighborhood centered on this pixel, the gray-scale values of the elevator door image are divided into L levels, a two-dimensional histogram N(i,j) is preset, i = f(x,y), j = h(x,y), f ij represents the frequency of occurrence of (i,j), and its corresponding joint probability density p ij The expression of is where 0≤i,j≤L-1, which improves the accuracy of image recognition.
[0052] Optionally, the contrast between the foreground color and the background color of the detection identification point is strong. According to the gray-scale characteristics of the marked points, the contour after edge detection is screened. The central area of the preset detection identification point is the foreground color E 1 , the gray-scale value of the pixel point is close to white, and M f is used to represent E 1 The average value of the pixels within the range; the annular area is the background color E 2 , the gray-scale value of the pixel point is close to black, and M b is used to represent E 2 The average value of the pixels within the range, then M f , M b should satisfy where M f represents the threshold for distinguishing the foreground color and the background color of the image, and ΔM t represents the minimum threshold that the difference between the foreground color and the background color should satisfy.
[0053] In this embodiment, the face recognition unit includes a model selection module. The model selection module uses the recognition signal and the verification signal, and the total loss is the weighted sum of the recognition signal and the verification signal loss. The weights are dynamically adjusted during the training process. The expressions of their respective loss functions are where \(f\) represents the feature vector, \(\theta\) represents the network parameters, \(t\) represents the target class, \(p\) represents the target probability distribution, represents the predicted probability distribution. When \(i = 1\), \(p\) i = 1. When the two samples are of the same person, \(y\) is 1, otherwise -1. The expression for training using the optimizer is where \(m\) t , \(v\) t represent the first and second order momenta, \(\beta\) 1 , \(\beta\) 2 represent the momentum parameters, represents the correction value, \(W\) t represents the parameters of the network at time \(t\).
[0054] It should be noted that the face recognition unit further includes a model evaluation module. The model evaluation module uses accuracy, precision, and recall as evaluation indicators for the model generalization ability. The authenticity of the model prediction result is determined by a confidence threshold, and the confidence threshold is the IoU threshold. The IoU threshold is set to 0.5. If the confidence is greater than 0.5, it is true; if the confidence is less than 0.5, it is false. Accuracy: Precision: Recall: The execution process of the face detection unit includes: detecting a face and extracting the feature vector of the face; comparing the feature vector with all the feature information in the label to find the feature vector with the closest Euclidean distance. Using the distance between these two vectors as the key value and the identity information corresponding to the corresponding feature vector in the label as the value, pair-wise save them to the temporary dictionary of the database unit; then, according to the number of detected people, preset as \(n\), extract the values corresponding to the first \(n\) values with the smallest key values in the temporary dictionary as the recognized information, and send out the command signal for setting the floor to complete the elevator touchless control.
[0055] Referring to Figure 2 , the present invention also provides an elevator touchless control method, including the following steps:
[0056] S1: Obtain a video stream from a binocular camera, obtain image information from the video stream, and input the image information into the face detection unit and the face recognition unit for processing to obtain a face image. Perform binocular ranging on the faces in the video according to the face image;
[0057] S2: Detect the face from the face image to obtain the position and size of the face. Among them, a deep learning detection algorithm is used for face detection;
[0058] S3: Extract the features of the face in the detected face image to obtain face features. Among them, through convolution, activation, and pooling operations, the feature vector corresponding to the face is extracted, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed the preset threshold, then there is no such person in the database. If it exceeds the preset threshold, the face identity information corresponding to the maximum value is taken;
[0059] S4: Display the face recognition result. The recognized face is marked with a square box, and the identity information and the state of the elevator are given. Among them, the elevator state includes the opening and closing of the elevator door, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state;
[0060] S5: Store the face feature vector, the identity information corresponding to the face feature vector, and the state of the elevator.
[0061] In this embodiment, the preprocessing of the picture includes separately processing the left and right parts of the picture frame, and adjusting the picture to the input size of the face detection network. After the processing is completed, only the picture frame of the left video is input into the network for detection. If a face is detected, a frame is drawn with the detected coordinate points in the output. If no face is detected, it is judged whether the current elevator state is 011. If so, it means that the almond has entered the elevator at this time and the floor information has been set, and then an instruction is issued to close the elevator door and start running the elevator; if not, it means that the elevator is not waiting for use on the first floor and no user who needs to take the elevator is found on the first floor. At this time, the video stream is continuously read for processing. When a face is detected, the abscissas of the upper left corner and the lower right corner of the face frame are extracted, and the arithmetic mean is performed respectively to obtain the center point of the face frame. Face detection is performed on the picture frame of the right camera, and the center point coordinates of the corresponding face frame of the right camera are found in the same way. According to the coordinates of the corresponding points on the left and right cameras, the coordinates of this person in space can be obtained by substituting them into the binocular ranging principle formula, so as to obtain the distance. It is possible to perform ranging on the detected face and quickly recognize the face, thereby improving the face recognition rate and accuracy of the elevator, and also improving the user experience.
[0062] In all the examples shown and described here, any specific value should be construed as merely exemplary, not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0063] It should be noted that: Similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0064] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. A sensorless elevator control system, It is characterized in that include: The binocular vision unit is used to obtain a video stream from a binocular camera, obtain image information from the video stream and input the image information into a face detection unit and a face recognition unit for processing to obtain a face image, and perform binocular distance measurement on the face in the video according to the face image; A face detection unit, used to detect a face from a face image and obtain the position and size of the face, wherein a deep learning detection algorithm is used for face detection; A face recognition unit is used to extract features of the face in the detected face image to obtain face features, wherein a feature vector corresponding to the face is extracted through convolution, activation and pooling operations, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed a preset threshold, the person does not exist in the database. If the matching degree exceeds the preset threshold, the face identity information corresponding to the maximum value is taken; A display unit is used to display the face recognition result, the recognized face is marked with a box, and the identity information and the state of the elevator are given, wherein the elevator state includes whether the elevator door is open or closed, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state; A database unit, used to store a facial feature vector, identity information corresponding to the facial feature vector, and a state of the elevator; The display unit includes an image preprocessing module for processing the elevator door image. The image preprocessing module uses the Gaussian filtering algorithm to perform image smoothing on the elevator door image to obtain the smoothed image g(x, y), and the corresponding expression is g 1 (x, y) = g(x, y) * G(x, y), where σ represents the standard variance of Gaussian filtering; Taking a 3×3 template as an example, within the template range, calculate the gray - level difference grd(x, y) between the current pixel point g(x, y) and other pixels. The corresponding expression is Let the Gaussian smoothing factor α = 1 - grd(x, y) / 255, and smooth the image through the Gaussian smoothing factor to obtain the improved smoothed image g 2 The expression of g(x, y) is 2 g(x, y) = αg(x, y)+(1 - α)g(x, y). When the image pixels are in the edge region, the greater the gray difference grd(x, y) between the pixels, the smaller the Gaussian smoothing factor α. 1 (x,y)+(1-α)g(x,y), when the image pixels are in the edge region, the greater the gray difference grd(x,y) between the pixels, the smaller the Gaussian smoothing factor α.
2. The elevator sensorless control system according to claim 1, It is characterized in that The display unit further includes an edge detection module for processing the elevator door image. The edge detection module adopts an adaptive Canny edge detection algorithm and calculates the smoothed image g by solving the finite difference mean 2 The expressions for the gradient magnitude and gradient direction of g where P x [x, y] represents the partial derivative function in the x direction, and P y [x, y] represents the partial derivative function in the y direction. The gradient magnitude is The gradient direction is θ[x, y] = arctan(P y [x, y] / P x [x, y]). By adding the gradient magnitudes in the 45° direction and 135° direction, the expression for the total gradient magnitude is where P x [x, y], P y [x, y], P xy [x, y], and P yx [x, y] represent the partial derivative functions in the x-axis direction, y-axis direction, 45° direction, and 135° direction, respectively; The Canny operator detects edges by setting high and low thresholds. If the gradient amplitude of a pixel is higher than the high threshold, the pixel is an edge point. If the gradient amplitude of a pixel is lower than the low threshold, the pixel is not an edge point.
3. The elevator sensorless control system according to claim 2, It is characterized in that The display unit also includes an edge screening module. The edge screening module uses image representation to extract detection mark point information. The detection mark points include coded mark points and non-coded mark points. The coded mark points and non-coded mark points are circles of the same size and are elliptical after CCD imaging. The execution process of the edge screening module includes: The ellipse contour is a convex closed contour. The distance between the starting pixel and the ending pixel of the curve segment after edge tracing is compared. If the distance is less than the threshold, the contour is determined to be closed, otherwise it is noise and is removed. Select a projection angle, that is, the angle between the normal direction of the plane where the identification point is located and the projection direction is between 0° and 60°, and take a picture of the elevator door. Within this angle range, the contour perimeter L and the elliptical area S of the identification point in the elevator door image respectively satisfy where L min and L max respectively represent the minimum and maximum ranges that the contour perimeter of the identification point should satisfy within this angle range, and S min and S max respectively represent the minimum and maximum ranges that the contour area of the identification point should satisfy within this angle range; The geometric properties of the ellipse will be utilized to screen the target edge contour, and the ellipse contours that meet the requirements within the range of 0°-60° will be screened. The expressions for screening the contour by the aspect ratio of the major and minor axes of the ellipse, the shape factor R, and the roundness C of the ellipse are where a and b respectively represent the major and minor axes of the ellipse, and are required to be less than the threshold M.
4. The elevator sensorless control system according to claim 3, It is characterized in that Also includes: The contrast between the foreground color and the background color of the detection identification point is strong. The contours after edge detection are filtered according to the gray-scale characteristics of the marked points. The center area of the preset detection identification point is the foreground color E 1 , the gray value of the pixel point is close to white, and is represented by M f for E 1 the average value of the pixels within the range; the annular area is the background color E 2 , the gray value of the pixel point is close to black, and is represented by M b for E 2 the average value of the pixels within the range, then M f , M b should satisfy where M f represents the threshold for distinguishing the foreground color and the background color of the image, and ΔM t represents the minimum threshold that the difference between the foreground color and the background color should satisfy.
5. The elevator sensorless control system according to claim 1, It is characterized in that The face recognition unit includes a model selection module. The model selection module uses recognition signals and verification signals. The total loss is the weighted sum of the recognition signal loss and the verification signal loss. The weights are dynamically adjusted during the training process. The expressions of their respective loss functions are where f represents the feature vector, θ represents the network parameters, t represents the target category, p represents the target probability distribution, represents the predicted probability distribution. When i = 1, p i = 1, y = 1 when the two samples are of the same person, otherwise y = -1; The expression for training using the optimizer is where m t , v t represent the first and second order momenta, and β 1 , β 2 represent the momentum parameters, represents the correction value, and W t represents the parameters of the network at time t.
6. The elevator sensorless control system according to claim 5, It is characterized in that The face recognition unit also includes a model evaluation module. The model evaluation module uses accuracy, precision, and recall as evaluation indicators for the model's generalization ability. The true or false of the model prediction result is determined by a confidence threshold, and the confidence threshold is the IoU threshold. The IoU threshold is set to 0.
5. If the confidence is greater than 0.5, it is true; if the confidence is less than 0.5, it is false. Accuracy: Precision: Recall:
7. The elevator sensorless control system according to claim 1, It is characterized in that The execution process of the face detection unit includes: Detect a face and extract the feature vector of the face; Compare the feature vector with all the feature information in the label, find the feature vector with the closest Euclidean distance, use the distance between the two vectors as the key value, and the identity information corresponding to the corresponding feature vector in the label as the value, and save them in pairs to the temporary dictionary of the database unit; Preset it as n according to the detected number of people, extract the values of the first n values with the smallest key values in the temporary dictionary as the recognized information, and send out the instruction signal for setting the floor to complete the elevator touchless control.
8. The elevator touchless control system according to claim 1, characterized in that The binocular vision unit includes a calibration module. The calibration module obtains the position information between the binocular cameras through binocular calibration. The positional relationship can be described by a translation matrix P and a rotation matrix T. By solving the translation and rotation matrices between the left and right camera coordinates and the world coordinate, the expression for the positional relationship between the left and right camera coordinates can be obtained as where P l represents the coordinates of P in the left camera coordinate system in the world coordinate system. Through the rotation matrix R l and the translation matrix T l the conversion between the left camera coordinate and the world coordinate is realized; P r represents the coordinates of P in the right camera coordinate system in the world coordinate system. Through the rotation matrix P r and the translation matrix T r the conversion between the right camera coordinate and the world coordinate is realized.
9. An elevator touchless control method for the elevator touchless control system according to any one of claims 1-8, characterized in that it includes the following steps: Obtain a video stream from a binocular camera, obtain image information from the video stream, input the image information into a face detection unit and a face recognition unit for processing to obtain a face image, and perform binocular ranging on the faces in the video according to the face image; Detect the face from the face image to obtain the position and size of the face. Among them, a deep learning detection algorithm is used for face detection; Extract the face features of the face in the detected face image to obtain face features. Among them, the feature vector corresponding to the face is extracted through convolution, activation and pooling operations, and the feature vector is compared one by one with the face feature vectors summarized in the database. If the matching degree of the face in the database does not exceed the preset threshold, there is no such person in the database. If it exceeds the preset threshold, the face identity information corresponding to the maximum value is taken; Display the face recognition result, mark the recognized face with a square box, and give the identity information and the state of the elevator. Among them, the elevator state includes the opening and closing of the elevator door, whether the elevator is in a moving state or a stopped state, and the floor corresponding to the stopped state; Store the face feature vector, the identity information corresponding to the face feature vector, and the state of the elevator.
Citation Information
Patent Citations
Optical fiber link component identification method and device, equipment and storage medium
CN113591787A
Method and system for carrying out enhancement processing on image edge and medium
CN114529459A