A virtual elevator button method

Through the dual-cascade convolutional neural network, real-time detection and tracking of user fingertips and identifying interactive gestures, the contactlessness of traditional elevator key operation is achieved, solving the problems of space limitations and infection risks of traditional elevator operation, and improving user interaction experience and system reliability.

CN115494939BActive Publication Date: 2025-05-23SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210930685.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-23
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The physical button system of traditional elevators has limitations in operating space, large time occupies and is susceptible to dust and oil pollution, which increases the risk of cross-infection and physiological safety hazards of users.

Method used

A dual-cascade convolutional neural network is used to obtain user images through the camera, detect and track user fingertips in real time, identify three interactive gestures: mouse, key and end, and realize non-contact elevator key operation.

Benefits of technology

It improves the friendship and nature of the user's interaction with elevator control, enhances the reliability and cleanliness of the system, reduces user operation time, and reduces the risk of cross-infection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115494939B_ABST
    Figure CN115494939B_ABST
Patent Text Reader

Abstract

The present invention discloses a virtual elevator button method, which includes: an image acquisition step, in which a user image is acquired through a camera; a fingertip detection step, in which a hand selection box and fingertip coordinates in the user image are acquired in real time, thereby realizing user fingertip positioning and tracking; a gesture recognition step, in which three interactive gestures of the user's fingertips are sequentially recognized, namely, mouse, button, and end, thereby forming an elevator control instruction; and an elevator control step, in which non-contact button control instructions are realized according to the elevator control instruction. The present invention realizes non-contact button operation of the elevator through the image acquisition step, the fingertip detection step, the gesture recognition step, and the elevator control step, refines the elevator control model, improves the elevator operation speed and utilization rate, thereby optimizing the elevator button controllability and flexibility, which can effectively save user time and improve user experience and comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of virtual key technology and intelligent elevator control technology, and in particular to a virtual elevator key method. Background Art

[0002] With the development of modern society and science and technology, people live and work on higher and higher floors. Elevators have gradually become an integral part of high-rise buildings and an important device for people's daily life, work and travel. However, most elevators still continue to operate in the traditional way. In today's era of intelligent tide, the intelligence of elevators is particularly important.

[0003] During the peak hours of daily office work, the population is dense and the amount of personnel activities is large, so there is a great demand for elevator use. Traditional vertical elevators require users to physically contact the elevator buttons, which leads to limited user control space, and physical contact can become a path for the spread of germs, increasing the risk of cross-infection and physiological safety hazards for users; therefore, improving the efficiency of elevator utilization through intelligent means and refining the elevator model are what society needs.

[0004] Traditional elevators rely on physical button systems to trigger operations. The elevator system triggers a long time and the operating space is small, which takes up users' valuable time to a certain extent. Existing non-contact elevator button systems such as infrared and capacitance are easily affected by dust, oil, etc. and are not sensitive, and require regular maintenance. Summary of the invention

[0005] The purpose of the present invention is to solve the above defects in the prior art and provide a virtual elevator key method. The method realizes elevator key operation without physical mechanical keys, and improves the friendliness and naturalness of the interaction between the user and the elevator control through intelligent human-computer interaction.

[0006] The purpose of the present invention can be achieved by adopting the following technical solutions:

[0007] A virtual elevator key method, the virtual elevator key method comprising:

[0008] An image acquisition step is to acquire a user image through a camera;

[0009] The fingertip detection step is to obtain the hand selection box and fingertip coordinates in the scene image in real time through a dual cascade convolutional neural network, and locate and track the user's fingertips. At the same time, the fingertip position and virtual buttons are displayed in real time through a display; wherein the dual cascade convolutional neural network includes a cascaded hand detection network and a multi-fingertip detection network, the hand detection network is used to locate the bounding box of the user's hand area in the scene image, and normalize the output of the upper left corner and lower right corner coordinates of the bounding box; the multi-fingertip detection network locates and tracks the user's fingertips by performing image processing, feature extraction and coordinate regression on the located hand area;

[0010] The gesture recognition step recognizes the three interactive gestures of the user's fingertips, namely, mouse, key and end, and then forms the elevator control instructions;

[0011] Elevator control steps: non-contact button operation according to elevator control instructions.

[0012] Furthermore, the working process of the hand detection network is as follows:

[0013] Ego-Gesture contains various gesture images in first-person perspective and complex environments. The Ego-Gesture dataset is used for training so that the network can better adapt to the application scenarios of elevator cars. In order to more accurately identify the hand area, a convolutional neural network is used as a feature extractor to generate a hand feature map;

[0014] Using the fully connected layer as a regressor, the coordinates of the upper left and lower right corners of the bounding box of the hand area are located using the feature regression in the hand feature map;

[0015] The hand detection network adopts gradient back propagation to perform model training. The first total loss function in the training process includes a first coordinate regression loss function, a first confidence loss function and a category loss function, so that the network can more accurately output a bounding box containing the hand area.

[0016] The first coordinate regression loss function calculates the width and height error of the bounding box, as shown in formula (1):

[0017]

[0018] In the formula, x i ,y i 、w i 、h i are the true values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the i-th bounding box obtained for the images trained in the same batch; are the predicted values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the corresponding i-th bounding box, respectively; λ and α are the first and second constant coefficients used to balance the magnitude and convergence speed of the two parts of the loss;

[0019] The first confidence loss function calculates the error of the probability that the bounding box contains objects of various categories, as shown in formula (2):

[0020]

[0021] In the formula, c i The true confidence value of the i-th positive sample box obtained for the same batch of training images, is the confidence prediction value of the corresponding i-th positive sample box; c j The confidence value of the j-th background box obtained for the same batch of training images, is the confidence prediction value of the corresponding j-th background box;

[0022] The category loss function calculates the error of object classification, as shown in formula (3):

[0023]

[0024] In the formula, f i is the true value predicted for the i-th image category in the same batch of training, Output the predicted value of the i-th image category corresponding to the double cascade convolutional neural network.

[0025] Furthermore, the working process of the multi-finger tip detection network is as follows:

[0026] Performing a scale transformation on the hand region image in the bounding box detected by the hand detection network;

[0027] Using convolutional neural network as feature extractor to generate multi-keypoint hand feature map;

[0028] A multi-branched fully connected layer is used as a regressor to complete the fingertip localization task and the fingertip confidence prediction task, the multi-keypoint hand feature map is used as input, the fingertip position is predicted through the fingertip localization task, and the fingertip confidence prediction task is used to predict the probability of the existence of the fingertip;

[0029] The second total loss function in the multi-fingertip detection network training process includes a second coordinate regression loss function and a second confidence loss function, wherein the second coordinate regression loss function, as shown in formula (4), describes the error between the predicted fingertip coordinates and the true coordinates:

[0030]

[0031] Among them, b is the number of samples in the same batch of training data, p i and q i is the true value of the fingertip coordinates located in the i-th hand region, and The predicted value of the fingertip coordinates located in the i-th hand region;

[0032] The second confidence loss function describes the error in predicting the probability of fingertip existence, as shown in formula (5):

[0033]

[0034] Among them, u i is the true value of the fingertip confidence of the i-th hand region positioning, The predicted value of the fingertip confidence for the localization of the i-th hand region;

[0035] The second total loss function is shown in formula (6):

[0036] Loss = ξ·loss p +β·loss u (6)

[0037] Among them, ξ and β are the first and second model hyperparameters used to balance the contribution of the two branches to the loss during the optimization process.

[0038] Furthermore, the multi-finger tip detection network uses the fingertip coordinates and their confidence as outputs, wherein the fingertip coordinates are the multi-channel output mean, and the calculation formula is shown in formula (7):

[0039]

[0040] Among them, y k is the output of a single channel of the double-cascaded convolutional neural network, C is the number of channels, and the final prediction result of the fingertip coordinates will be constrained by the confidence output. Only when the confidence prediction value is higher than the first threshold th 1 The fingertip is considered to be real, the predicted coordinates are output normally, and the confidence prediction value is lower than the first threshold th 1 The fingertip is considered non-existent and the predicted coordinates are blocked.

[0041] Furthermore, the gesture recognition step process is as follows:

[0042] The results of multi-finger tip detection correspond to different gesture categories, based on which finger tip encoding is performed to construct a gesture encoding library;

[0043] Fingertip encoding is performed using fingertip location coordinates without using a separate classifier;

[0044] By matching the fingertip code with the fingertip code in the code library, the corresponding gesture type is obtained, thereby achieving the purpose of gesture recognition.

[0045] Furthermore, in order to complete the elevator control process more concisely, the gesture types include three interactive gestures: mouse, button and end. Among them, the mouse gesture realizes the wake-up and fingertip tracking functions of the virtual elevator button function; the button gesture realizes the trigger button function; and the end gesture realizes the termination of the button control function.

[0046] Furthermore, in order to improve the efficiency of elevator operation, the awakening of the virtual elevator button function refers to the awakening of the virtual elevator button function when the system detects that the number of consecutive mouse gesture image frames reaches a second threshold value th. 2 When required, determine whether the user performs a wake-up operation to avoid accidental touches caused by direct key operations;

[0047] The trigger button function includes a debounce mechanism, which continuously scans the gesture state and when the number of frames maintained by the same gesture reaches a third threshold value th 3 When the user moves his hand, the corresponding operation is triggered to prevent the user from accidentally touching the elevator button;

[0048] Furthermore, the elevator control process is as follows:

[0049] When the gesture recognition step recognizes a mouse gesture and passes the wake-up mechanism to prevent false triggering, the virtual elevator button operation is awakened and the fingertip tracking function is turned on;

[0050] When the gesture recognition step recognizes a key gesture, the key is triggered according to the proximity principle;

[0051] When the gesture recognition step recognizes the end gesture, fingertip tracking and key detection are continued; when the end gesture is continuously recognized, key detection is terminated, the control unit of the elevator is triggered, and the elevator operation is controlled according to the pressed key, and the coordinate record is cleared.

[0052] Furthermore, the fingertip tracking is implemented by a fingertip tracking mechanism. The fingertip tracking mechanism records the coordinates of the fingertip movement trajectory in a time series. In order to smooth the movement of the fingertip, a linear filter is used to make the trajectory smoother within a certain length interval. The smooth trajectory path processing is shown in formulas (8) and (9):

[0053]

[0054] Where l is the smoothing length, p is the end point of the smoothing trajectory, and m i and n i are the horizontal and vertical coordinates of the i-th point in the coordinate record respectively.

[0055] Furthermore, the display screen is installed inside the elevator to display the user's fingertip position and the changes and current status of the elevator buttons in real time. When the elevator button is not triggered, it shows a first distinguishing color, which can be a light color, such as white; when triggered, it shows a second distinguishing color, which can be a dark color, such as red; after reaching the designated floor, the corresponding button color is restored.

[0056] Compared with the prior art, the present invention has the following advantages and effects:

[0057] (1) A dual-cascade neural network is used to first locate the user's hand and then track the fingertips, making the fingertip detection results more accurate. At the same time, an anti-false touch wake-up mechanism is added to ensure the reliability of the system.

[0058] (2) Through visual detection of the user's fingertips and recognition of gestures, non-contact key triggering is achieved, which is clean and hygienic. At the same time, the equipment is simple, the system performance is stable, and it is not easily affected by environmental factors, and has a long service life;

[0059] (3) The algorithm responds quickly, making the elevator operate in real time and effectively. At the same time, the present invention can realize simultaneous control by multiple people, so that user operation is not limited by space limitations, key control is more flexible, and friendly and natural human-computer interaction is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0061] Figure 1 It is a flow chart of a virtual elevator key pressing method of the present invention;

[0062] Figure 2 This is a diagram of the network structure of multi-fingertip detection in the dual-cascade convolutional neural network of the present invention. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0064] Example 1

[0065] like Figure 1 As shown, it is a flowchart of a virtual elevator button method provided by the first embodiment of the present invention, comprising the following steps:

[0066] An image acquisition step is to acquire an image inside the elevator car through a camera inside the elevator car;

[0067] In the fingertip detection step, the hand selection frame and fingertip coordinates in the scene image are obtained in real time through the trained dual-cascade convolutional neural network, so as to locate and track the user's fingertips. At the same time, the fingertip position and virtual buttons are displayed in real time on the display.

[0068] The dual cascade convolutional neural network includes a cascaded hand detection network and a multi-fingertip detection network. The hand detection network uses a YOLOV3 network. The image in the elevator car is input, and the network normalizes and outputs the coordinates of the upper left corner and lower right corner of the user's hand boundary box.

[0069] The hand detection network uses gradient back propagation to optimize the first total loss function for model training. The first total loss function includes the first coordinate regression loss function, the first confidence loss function, and the category loss function. The bounding box error, the error of the probability that the bounding box contains objects of different categories, and the error of object classification are calculated respectively. The smaller the value of the first total loss function, the more accurate the output bounding box containing the hand.

[0070] The first coordinate regression loss function is shown in formula (1):

[0071]

[0072] In the formula, x i ,y i 、w i 、h i are the true values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the i-th bounding box obtained for the images trained in the same batch; are the predicted values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the corresponding i-th bounding box, respectively; λ and α are the first and second constant coefficients used to balance the magnitude and convergence speed of the two parts of the loss;

[0073] Among them, the first confidence loss function is shown in formula (2):

[0074]

[0075] In the formula, c i The true confidence value of the i-th positive sample box obtained for the same batch of training images, is the confidence prediction value of the corresponding i-th positive sample box; c j The confidence value of the j-th background box obtained for the same batch of training images, is the confidence prediction value of the corresponding j-th background box;

[0076] Among them, the category loss function is shown in formula (3):

[0077]

[0078] In the formula, f i is the true value predicted for the i-th image category in the same batch of training, Output the predicted value of the i-th image category corresponding to the double cascade convolutional neural network.

[0079] The area image within the hand boundary box is cropped to a size of 128×128 and input into the multi-fingertip detection network. A convolutional neural network is used to extract features of the located hand area, and a hand feature map is used to locate fingertips and predict fingertip confidence.

[0080] For the fingertip localization task, the input feature map goes through two convolutional layers, a batch normalization layer, and an activation function, such as Figure 2 As shown, after upsampling and convolutional layers, the fingertip coordinates are finally averaged and output; the second coordinate regression loss function is set to describe the error between the predicted fingertip coordinates and the true coordinates, as shown in formula (4):

[0081]

[0082] Among them, b is the number of samples in the same batch of training data, p i and q i is the true value of the fingertip coordinates located in the i-th hand region, and The predicted value of the fingertip coordinates located in the i-th hand region;

[0083] For the fingertip confidence prediction task, after two linear layers, batch normalization layer, activation function such as Figure 2 As shown, after a linear layer and softmax, the fingertip confidence prediction value is output; the second confidence loss function is set to describe the error of the predicted fingertip existence probability, as shown in formula (5):

[0084]

[0085] Among them, u i is the true value of the fingertip confidence of the i-th hand region positioning, The predicted value of the fingertip confidence for the localization of the i-th hand region;

[0086] Combining the two tasks, the second total loss function is defined as shown in formula (6):

[0087] Loss = ξ·loss p +β·loss u (6)

[0088] Among them, ξ and β are the first and second model hyperparameters used to balance the contribution of the two branches to the loss during the optimization process. The second total loss function is minimized through reverse gradient propagation, and the final output confidence prediction value is higher than the first threshold th 1 The fingertip coordinates are the multi-channel output mean values, and the calculation formula is shown in formula (7):

[0089]

[0090] Among them, y k is the output of a single channel of the double cascade convolutional neural network, and C is the number of channels.

[0091] Gesture recognition step: Encode the detected coordinates of the user's fingertips and compare them with the encoding library, identify the three interactive gestures of the user's fingertips in turn: mouse, key and end, and then form the elevator control command. Specifically, it can be set that extending 3 fingers is the mouse, extending 1 finger is the key, and extending 5 fingers is the end;

[0092] The elevator control step is to implement non-contact key operation according to the elevator control command. Specifically, the user first extends three fingers and maintains them. When the number of frames exceeds the second threshold th 2 Requirements: Detect mouse gestures and enable fingertip tracking. Then, the user extends one finger to the button to be triggered and holds it. When the frame rate exceeds the third threshold th 3 When required, the button is triggered according to the principle of proximity; finally, the user extends 5 fingers, and when the system continuously recognizes the end gesture, the button detection is terminated, the elevator control unit is triggered, and the elevator operation is controlled according to the pressed button, and the coordinate record is cleared;

[0093] Example 2

[0094] This embodiment continues to disclose a virtual elevator button method, and the specific process is as follows:

[0095] An image acquisition step is to acquire an image inside the elevator car through a camera inside the elevator car;

[0096] The fingertip detection step uses a trained dual-cascade convolutional neural network to obtain the hand selection box and fingertip coordinates in the scene image in real time, including a hand detection network and a multi-fingertip detection network. This embodiment continues to use the trained YOLOV3 as the hand detection network, and crops the detected user hand area image into an image of 128×128 size, which is input into the multi-fingertip detection network;

[0097] The multi-finger tip detection network architecture disclosed in the present invention is as follows: Figure 2As shown, this embodiment uses the Ego-Gesture dataset for training, which contains various gesture images in the first-person perspective and complex environments. The input dataset is processed by cropping, mirroring and rotating data, and then feature extraction is performed to obtain image features of size 4×4×512, which are respectively input into two branches: the fingertip positioning branch and the confidence prediction branch, to obtain the fingertip coordinates and the fingertip confidence prediction respectively. In this way, the user's extended finger and position can be determined, and then gesture recognition can be performed.

[0098] The average detection error of the trained multi-finger prediction network compared with other advanced methods in fingertip detection accuracy is reduced by about 0.6 pixels, as shown in Table 1:

[0099] Table 1. Comparison of fingertip detection errors of the fingertip detection method in Example 2 and other methods

[0100] method Gesture "1" Gesture "2" Gesture "3" Gesture "4" Gesture "5" average FD 7.12 7.08 7.08 7.43 7.62 7.26 YOLSE 4.13 3.17 2.75 3.33 3.21 3.31 The present invention 3.76 2.66 2.55 2.52 2.50 2.79

[0101] The trained multi-finger detection network can run at an optimal speed of 5ms per frame on the GPU. The comparison results with other methods are shown in Table 2:

[0102] Table 2. Comparison of running speed of fingertip detection method in Example 2 and other methods on GPU

[0103] method Time required for each frame (ms) HeatmapFushionNet 42 YOLSE 15 Ours 5

[0104] The gesture recognition step sequentially recognizes the three interactive gestures of the user's fingertips, namely, mouse, key and end, and then forms an elevator control instruction. Specifically, it can be set that extending three fingers is a mouse, extending one finger is a key, and extending five fingers is an end;

[0105] Elevator control steps: non-contact button operation according to elevator control instructions.

[0106] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A virtual elevator button method, It is characterized in that The virtual elevator button method comprises: An image acquisition step is to acquire a user image through a camera; The fingertip detection step is to obtain the hand selection box and fingertip coordinates in the scene image in real time through a dual cascade convolutional neural network, and locate and track the user's fingertips. At the same time, the fingertip position and virtual buttons are displayed in real time through a display; wherein the dual cascade convolutional neural network includes a cascaded hand detection network and a multi-fingertip detection network, the hand detection network is used to locate the bounding box of the user's hand area in the scene image, and normalize the output of the upper left corner and lower right corner coordinates of the bounding box; the multi-fingertip detection network locates and tracks the user's fingertips by performing image processing, feature extraction and coordinate regression on the located hand area; The gesture recognition step recognizes the three interactive gestures of the user's fingertips, namely, mouse, key and end, and then forms the elevator control instructions; Elevator control steps: non-contact key operation according to elevator control instructions; The working process of the hand detection network is as follows: The various gesture images in the Ego-Gesture dataset are used as input, and a convolutional neural network is used as a feature extractor to generate hand feature maps. The Ego-Gesture dataset contains various gesture images in first-person perspective and complex environments. Use the fully connected layer as a regressor to regress the features in the hand feature map to locate the upper left and lower right corner coordinates of the bounding box of the hand area; The hand detection network uses a gradient back propagation method to perform model training. The first total loss function in the training process includes a first coordinate regression loss function, a first confidence loss function, and a category loss function. The coordinate regression loss function is shown in formula (1): In the formula, x i ,y i 、w i 、h i are the true values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the i-th bounding box obtained for the images trained in the same batch; are the predicted values ​​of the horizontal coordinate, vertical coordinate, length and width of the upper left corner of the corresponding i-th bounding box, respectively; λ and α are the first and second constant coefficients used to balance the magnitude and convergence speed of the two parts of the loss; The first confidence loss function is as shown in formula (2): In the formula, c i The true confidence value of the i-th positive sample box obtained for the same batch of training images, is the confidence prediction value of the corresponding i-th positive sample box; c j The confidence value of the j-th background box obtained for the same batch of training images, is the confidence prediction value of the corresponding j-th background box; The category loss function is shown in formula (3): In the formula, f i is the true value predicted for the i-th image category in the same batch of training, Output the predicted value of the i-th image category corresponding to the double cascade convolutional neural network; The working process of the multi-finger tip detection network is as follows: Performing a scale transformation on the hand region image in the bounding box detected by the hand detection network; Using convolutional neural network as feature extractor to generate multi-keypoint hand feature map; A multi-branched fully connected layer is used as a regressor to complete the fingertip localization task and the fingertip confidence prediction task, the multi-keypoint hand feature map is used as input, the fingertip position is predicted through the fingertip localization task, and the fingertip confidence prediction task is used to predict the probability of the existence of the fingertip; The second total loss function in the multi-finger detection network training process includes a second coordinate regression loss function and a second confidence loss function, wherein the second coordinate regression loss function is as shown in formula (4): Among them, b is the number of samples in the same batch of training data, p i and q i is the true value of the fingertip coordinates located in the i-th hand region, and The predicted value of the fingertip coordinates located in the i-th hand region; The second confidence loss function is as shown in formula (5): Among them, u i is the true value of the fingertip confidence of the i-th hand region positioning, The predicted value of the fingertip confidence for the localization of the i-th hand region; The second total loss function is shown in formula (6): Loss=ξ·loss p +β·loss u (6) Among them, ξ and β are the first and second model hyperparameters used to balance the contribution of the two branches to the loss during the optimization process.

2. A virtual elevator button method according to claim 1, It is characterized in that The multi-finger tip detection network uses the fingertip coordinates and their confidence as outputs, where the fingertip coordinates are the multi-channel output mean, and the calculation formula is shown in formula (7): Among them, y k is the output of a single channel of the double-cascaded convolutional neural network, C is the number of channels, and the final prediction result of the fingertip coordinates will be constrained by the confidence output. Only when the confidence prediction value is higher than the first threshold th 1 The fingertip is considered to be real, the predicted coordinates are output normally, and the confidence prediction value is lower than the first threshold th 1 The fingertip is considered non-existent and the predicted coordinates are blocked.

3. A virtual elevator button method according to claim 1, It is characterized in that The gesture recognition process is as follows: Encode different types of gestures with fingertips and construct a coding library; Fingertip encoding is performed using fingertip location coordinates without using a separate classifier; By matching the fingertip code with the fingertip code in the code library, the corresponding gesture type is obtained, thereby achieving the purpose of gesture recognition.

4. A virtual elevator button method according to claim 1, It is characterized in that The gesture types include three types of interactive gestures: mouse, key and end. Among them, the mouse gesture realizes the wake-up and fingertip tracking functions of the virtual elevator key function; the key gesture realizes the trigger key function; and the end gesture realizes the termination key control function.

5. A virtual elevator button method according to claim 4, It is characterized in that The awakening of the virtual elevator button function is realized based on the awakening mechanism of preventing false triggering. When the continuous mouse gestures reach the second threshold th 2 When required, determine whether the user performs a wake-up operation; The trigger button function includes a debounce mechanism, which continuously scans the gesture state and when the number of frames maintained by the same gesture reaches a third threshold value th 3 , triggers the corresponding action.

6. A virtual elevator button method according to claim 1, It is characterized in that The elevator control process is as follows: When the gesture recognition step recognizes a mouse gesture and passes the wake-up mechanism to prevent false triggering, the virtual elevator button operation is awakened and the fingertip tracking function is turned on; When the gesture recognition step recognizes a key gesture, the key is triggered according to the proximity principle; When the gesture recognition step recognizes the end gesture, fingertip tracking and key detection are continued; when the end gesture is continuously recognized, key detection is terminated, the control unit of the elevator is triggered, and the elevator operation is controlled according to the pressed key, and the coordinate record is cleared.

7. A virtual elevator button method according to claim 1, It is characterized in that The fingertip tracking is implemented by a fingertip tracking mechanism. The fingertip tracking mechanism records the coordinates of the fingertip movement trajectory in a time series and smoothes the trajectory path within a certain length interval. The smooth trajectory path processing is shown in formulas (8) and (9): Where l is the smoothing length, p is the end point of the smoothing trajectory, and m i and n i are the horizontal and vertical coordinates of the i-th point in the coordinate record respectively.

8. A virtual elevator button method according to claim 1, It is characterized in that The display is installed inside the elevator and is used to display the user's fingertip position and elevator buttons in real time. When the elevator button is not triggered, it shows the first distinguishing color; when triggered, it shows the second distinguishing color; after reaching the designated floor, the corresponding button color is restored.

Citation Information

Patent Citations

  • Egocentric vision in-the-air hand-writing and in-the-air interaction method based on cascade convolution nerve network

    CN105718878A

  • Application method of non-contact elevator control panel based on gesture control

    CN112114675A