Gesture recognition method and apparatus, and electronic device

By initializing the operating environment on the web page and dynamically selecting gesture recognition methods for deep learning frameworks, the problems of high costs and low universality in the existing technology are solved, and resource optimization and universality are achieved.

WO2025112938A1PCT designated stage expired Publication Date: 2025-06-05CHINA TELECOM CORP LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/124488
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-10-12
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing gesture recognition scheme based on deep learning has high costs in terms of data annotation and server resource consumption, and is not universal enough to cope with all user environments.

Method used

By initializing the running environment on the web page, the gesture recognition model in the two deep learning frameworks is run separately, and the optimal deep learning framework is selected for gesture recognition based on the resource consumption comparison results, realizing dynamic switching of the model and resource optimization.

Benefits of technology

Deploying the gesture recognition model on the web page reduces the consumption and cost of server resources, while improving the universality and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124488_05062025_PF_FP_ABST
    Figure CN2024124488_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a gesture recognition method and apparatus, and an electronic device. The method comprises: in response to a starting operation of a target object at a webpage end, initializing the current running environment; when the current running environment supports a target format, respectively running a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework on the basis of the current running environment, so as to obtain first resource consumption and second resource consumption; on the basis of a comparison result of the first resource consumption and the second resource consumption, determining a gesture recognition model in a target deep learning framework, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; and performing gesture recognition on the basis of the gesture recognition model in the target deep learning framework. The present application solves the technical problem in the relevant art of a significant amount of server resources and costs being consumed due to using a deep-learning-based gesture recognition solution to deploy a model at a cloud end.
Need to check novelty before this filing date? Find Prior Art

Description

Gesture recognition method, device and electronic device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 2023116077322 and application name “Method, device and electronic device for gesture recognition”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, and electronic device for gesture recognition. Background Art

[0003] Gesture recognition technology uses computer vision and artificial intelligence to analyze and identify human gestures. It primarily captures gesture information through input and output devices like cameras and sensors, then processes and analyzes it using algorithms to identify and interpret gestures.

[0004] Current gesture recognition technologies can be divided into two categories: sensor-based and vision-based. Sensor-based gesture recognition primarily utilizes sensors such as accelerometers and gyroscopes to capture gesture information and uses algorithms to analyze and recognize it. Vision-based gesture recognition, on the other hand, utilizes cameras to capture gesture information and uses computer vision and artificial intelligence algorithms to analyze and recognize it. Deep learning-based gesture recognition is the mainstream approach to vision-based gesture recognition. While deep learning-based gesture recognition has achieved impressive accuracy, several challenges remain.

[0005] 1. Deep learning, as a data-driven task, requires a large amount of labeled data to train models, and model accuracy is directly proportional to the amount of data. Training a high-performing model typically requires hundreds of thousands or even millions of labeled data points. Labeling data requires significant manpower and time, and purchasing data is also a significant expense.

[0006] 2. Deep learning, as a computationally intensive task, places significant demands on hardware resources. When using gesture recognition solutions based on deep learning, the model is typically deployed in the cloud. This requires a certain amount of server resources, which is directly proportional to the number of calls. This consumes a significant amount of server resources and costs.

[0007] 3. Current gesture recognition solutions based on deep learning can only provide a single model and cannot cope with all user environments. Its universality needs to be improved.

[0008] To address the above-mentioned problems, no effective solutions have been proposed so far.

[0009] Summary of the Invention

[0010] The embodiments of the present application provide a method, apparatus, and electronic device for gesture recognition to at least solve the technical problem in related technologies of using a deep learning-based gesture recognition solution, deploying the model in the cloud, and consuming a large amount of server resources and costs.

[0011] According to one aspect of an embodiment of the present application, a method for gesture recognition is provided, including: initializing a current operating environment in response to a startup operation of a target object on a web page; when the current operating environment supports a target format, respectively running a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operating environment to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; determining a gesture recognition model in a target deep learning framework based on a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; and performing gesture recognition according to the gesture recognition model in the target deep learning framework.

[0012] In some embodiments, the method further includes: after initializing the current operating environment, if the current operating environment does not support the target format, calling a gesture recognition model deployed in the server in the form of a backend service, wherein the gesture recognition model is used to recognize the gesture of the target object.

[0013] In some embodiments, the gesture recognition model in the target deep learning framework is determined based on the comparison result of the first resource consumption and the second resource consumption, including: when the first resource consumption is greater than the second resource consumption, the gesture recognition model in the second deep learning framework is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model in the first deep learning framework is released; when the first resource consumption is less than the second resource consumption, the gesture recognition model in the first deep learning framework is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model in the second deep learning framework is released; when the first resource consumption is equal to the second resource consumption, the gesture recognition model in the first deep learning framework or the second deep learning framework is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model outside the target deep learning framework is released.

[0014] In some embodiments, gesture recognition is performed based on a gesture recognition model in a target deep learning framework, including: recognizing the gesture of the target object at a preset frame rate to obtain a recognition result; when the accuracy of the target gesture recognized in the recognition result is greater than a preset threshold, triggering a special effect corresponding to the target gesture.

[0015] In some embodiments, a gesture recognition model is obtained by: inputting training data into a first original model and then training it through a momentum encoder to obtain a first feature extraction model, wherein the training data is unlabeled image data; cross-entropy training is performed on a second original model based on the first feature extraction model and the training data to obtain a second feature extraction model, wherein the second original model is a simplified model of the first original model; and the second feature extraction model is trained using labeled data to obtain a gesture recognition model, wherein the gesture recognition model includes a hand detection model and a gesture classification model.

[0016] In some embodiments, cross-entropy training is performed on the second original model based on the first feature extraction model and the training data to obtain a second feature extraction model, including: inputting the training data into the first feature extraction model and the second original model respectively to obtain a first output and a second output; storing the first output and the historical output of the first feature extraction model into an instance queue; performing an inner product operation on the output in the instance queue and the first output to obtain a first probability distribution; performing an inner product operation on the output in the instance queue and the second output to obtain a second probability distribution; performing a cross-entropy operation on the first probability distribution and the second probability distribution, and training the second original model based on the result of the cross-entropy operation to obtain a second feature extraction model.

[0017] In some embodiments, the second feature extraction model is trained using the labeled data to obtain a gesture recognition model, including: using the second feature extraction model as the backbone network, freezing the feature extraction layer of the second feature extraction model, and adding a multi-layer perceptron layer after the feature extraction layer to obtain a detection network; training the detection network based on the first category of data in the labeled data to obtain a hand detection model, wherein the hand detection model is used to identify whether it is a gesture; training the detection network based on the second category of data in the labeled data to obtain a gesture classification model, wherein the gesture classification model is used to identify the category of the gesture.

[0018] In some embodiments, the method also includes: after obtaining the gesture recognition model, converting the gesture recognition model into a gesture recognition model supported by the first deep learning framework according to the conversion tool of the first deep learning framework, to obtain the gesture recognition model in the first deep learning framework; converting the gesture recognition model into a gesture recognition model supported by the second deep learning framework according to the conversion tool of the second deep learning framework, to obtain the gesture recognition model in the second deep learning framework; using the target language to encapsulate the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework into a software development kit; using the format conversion tool to convert the code in the software development kit into the target format.

[0019] According to another aspect of an embodiment of the present application, a gesture recognition device is also provided, including: a processing module for initializing a current operating environment in response to a startup operation of a target object on a web page; an operating module for respectively operating a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operating environment, when the current operating environment supports a target format, to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; a determination module for determining the gesture recognition model in the target deep learning framework based on a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; and a recognition module for performing gesture recognition based on the gesture recognition model in the target deep learning framework.

[0020] According to another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory for storing program instructions; a processor, connected to the memory, for executing program instructions to implement the following functions: initializing the current operating environment in response to the startup operation of the target object on the web page; when the current operating environment supports the target format, running the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework respectively according to the current operating environment to obtain first resource consumption and second resource consumption, wherein the target format is a format supported by the web page; determining the gesture recognition model in the target deep learning framework based on the comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; performing gesture recognition according to the gesture recognition model in the target deep learning framework.

[0021] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned gesture recognition method by running the computer program.

[0022] In an embodiment of the present application, the current operating environment is initialized by responding to the startup operation of the target object on the web page; when the current operating environment supports the target format, the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework are respectively run according to the current operating environment to obtain the first resource consumption and the second resource consumption, wherein the target format is the format supported by the web page; based on the comparison result of the first resource consumption and the second resource consumption, the gesture recognition model in the target deep learning framework is determined, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; gesture recognition is performed according to the gesture recognition model in the target deep learning framework, thereby achieving the purpose of deploying the gesture recognition model on the web page, thereby realizing the technical effect of allocating the server cost to each user and saving the server overhead, thereby solving the technical problem in the related technology of using a deep learning-based gesture recognition solution to deploy the model in the cloud, consuming a large amount of server resources and costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0024] FIG1 is a hardware structure block diagram of a computer terminal for implementing a method for gesture recognition according to an embodiment of the present application;

[0025] FIG2 is a flowchart of a method for gesture recognition according to an embodiment of the present application;

[0026] FIG3 is an overall flow chart of a front-end gesture recognition system according to an embodiment of the present application;

[0027] FIG4 is a schematic diagram of a self-supervised distillation process according to an embodiment of the present application;

[0028] FIG5 is a flow chart of a process for building a front-end gesture recognition system based on self-supervised distillation according to an embodiment of the present application;

[0029] FIG6 is a structural diagram of a gesture recognition device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] First, some nouns or terms that appear in the process of explaining the embodiments of this application are subject to the following explanations:

[0033] Unlabeled dataset: refers to a dataset that does not provide clear annotations or labels. This means that each sample in the dataset lacks clear classification or label information.

[0034] Cross-entropy: A key concept in Shannon's information theory, it measures the difference between two probability distributions. Cross-entropy can be used as a loss function in neural networks (machine learning). Where p represents the distribution of true labels and q represents the distribution of predicted labels from the trained model, the cross-entropy loss function measures the similarity between p and q.

[0035] NanoDet: An ultra-fast and lightweight anchor-free object detection model for mobile devices.

[0036] Knowledge distillation: Knowledge distillation builds a lightweight, smaller model and uses the supervisory information from a larger, higher-performing model to train it, aiming for better performance and accuracy. This larger model is called the teacher, and the smaller model is called the student. The supervisory information output by the teacher model is called knowledge, and the process by which the student learns and transfers this supervisory information is called distillation.

[0037] wasm: WebAssembly. WebAssembly is a binary instruction set based on a stack-based virtual machine. It can be used as a compilation target for programming languages ​​and can be deployed in web client and server applications.

[0038] Emscripten: Emscripten is a compiler that can compile C / C++ code into JavaScript glue code. Emscripten can compile C / C++ code into code in the WebAssembly programming language.

[0039] NCNN: A high-performance neural network forward computing framework optimized for mobile devices.

[0040] Openvino: A toolkit for optimizing and deploying AI reasoning, primarily for optimizing deep reasoning.

[0041] JS (JavaScript): A directly interpreted scripting language.

[0042] MLP (Multilayer Perceptron): Multilayer Perceptron is also called Artificial Neural Network (ANN). In addition to the input and output layers, it can have multiple hidden layers.

[0043] MobileNetV2: A lightweight convolutional neural network.

[0044] Resnet101: A deep residual network.

[0045] The gesture recognition method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal for implementing the gesture recognition method. As shown in Figure 1, the computer terminal 10 may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. It will be understood by those skilled in the art that the structure shown in Figure 1 is merely illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than shown in Figure 1, or have a configuration different from that shown in Figure 1.

[0046] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0047] The memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the gesture recognition method in the embodiments of the present application. The processor executes the software programs and modules stored in the memory 104 to execute various functional applications and data processing, thereby implementing the gesture recognition method described above. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0048] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0049] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0050] It should be noted that, in some optional embodiments, the computer terminal shown in FIG1 may include hardware components (including circuits), software components (including computer code stored on a computer-readable medium), or a combination of hardware components and software components. It should be noted that FIG1 is only an example of a specific embodiment and is intended to illustrate the types of components that may be present in the computer terminal.

[0051] In the above operating environment, an embodiment of the present application provides an embodiment of a method for gesture recognition. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0052] FIG2 is a flow chart of a method for gesture recognition according to an embodiment of the present application. As shown in FIG2 , the method includes the following steps:

[0053] Step S202: Initialize the current operating environment in response to the start-up operation of the target object on the web page.

[0054] In the above step S202, the web page can be a page of the gesture recognition system. When the user opens the page, the gesture recognition system will perform an initialization operation to determine whether the current user environment (or current operating environment) supports wasm (webassembly, the target format below), and then enter different subsequent processes according to the judgment result.

[0055] Step S204: If the current operating environment supports the target format, the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework are respectively run according to the current operating environment to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page.

[0056] In the above step S204, if the current operating environment supports wasm, the system will perform two empty operations on the web page, that is, run the gesture recognition models in the ncnn (that is, the above-mentioned first deep learning framework) and openvino (that is, the above-mentioned second deep learning framework) versions respectively, and obtain the resource consumption of the two types of models, that is, the above-mentioned first resource consumption and second resource consumption.

[0057] Step S206 : determining a gesture recognition model in a target deep learning framework based on a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework.

[0058] In the above step S206, by comparing the first resource consumption and the second resource consumption, the gesture recognition model in the framework with the best performance in the current user environment (that is, the gesture recognition model in the above target deep learning framework) is selected according to the comparison result of the first resource consumption and the second resource consumption, and the gesture recognition models in the unselected frameworks are released.

[0059] Step S208 : performing gesture recognition based on a gesture recognition model in a target deep learning framework.

[0060] In the above step S208, when the user turns on the gesture recognition function on the web page, the gesture can be recognized according to the gesture recognition model in the selected target deep learning framework.

[0061] Through steps S202 to S208 above, the gesture recognition model is deployed on the web page, thereby allocating server costs to each user and reducing server overhead. This solves the technical problem of deploying the model in the cloud using deep learning-based gesture recognition solutions, which consumes significant server resources and costs. This is explained in detail below.

[0062] In step S202 of the gesture recognition method described above, after initializing the current operating environment, the method further includes the following steps: if the current operating environment does not support the target format, calling a gesture recognition model deployed in the server in the form of a backend service, wherein the gesture recognition model is used to recognize the gesture of the target object.

[0063] In the embodiment of the present application, after determining whether the current user environment supports the wasm format, different subsequent processes are entered according to the determination result. If the determination result is that the current operating environment supports wasm, the system will execute the process in step S204 above. If the determination result is that the current operating environment does not support wasm, the system will automatically switch to the backend service form and call the gesture recognition service through the http interface.

[0064] In step S206 of the gesture recognition method, determining the gesture recognition model in the target deep learning framework based on the comparison result of the first resource consumption and the second resource consumption specifically includes the following steps: when the first resource consumption is greater than the second resource consumption, using the gesture recognition model in the second deep learning framework as the gesture recognition model in the target deep learning framework, and releasing the gesture recognition model in the first deep learning framework; when the first resource consumption is less than the second resource consumption, using the gesture recognition model in the first deep learning framework as the gesture recognition model in the target deep learning framework, and releasing the gesture recognition model in the second deep learning framework; when the first resource consumption is equal to the second resource consumption, using the gesture recognition model in the first deep learning framework or the second deep learning framework as the gesture recognition model in the target deep learning framework, and releasing gesture recognition models outside the target deep learning framework.

[0065] In the embodiment of the present application, since it is necessary to select the gesture recognition model in the framework with the best performance in the current user environment as the gesture recognition model in the above-mentioned target deep learning framework, it is necessary to compare the first resource consumption and the second resource consumption. Specifically:

[0066] When the first resource consumption is less than the second resource consumption, it means that the gesture recognition model in the first deep learning framework (ncnn) corresponding to the first resource consumption is the most suitable for the current user environment performance. Therefore, the ncnn model is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model in the second deep learning framework, that is, the Openvino model, is released.

[0067] When the first resource consumption is greater than the second resource consumption, it means that the gesture recognition model in the second deep learning framework (Openvino) corresponding to the second resource consumption is the most suitable for the current user environment performance. Therefore, the Openvino model is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model in the first deep learning framework is released, that is, the ncnn model is released.

[0068] When the first resource consumption is equal to the second resource consumption, it means that both the ncnn model and the openvino model meet the performance of the current user environment. Therefore, one of the ncnn model and the openvino model is randomly selected as the gesture recognition model in the target deep learning framework, and the other model is released.

[0069] Through the above steps, the system automatically selects different links and models to maximize the universality of the system.

[0070] In step S208 of the above-mentioned gesture recognition method, gesture recognition is performed based on the gesture recognition model in the target deep learning framework, specifically including the following steps: recognizing the gesture of the target object at a preset frame rate to obtain a recognition result; when the accuracy of the target gesture recognized in the recognition result is greater than a preset threshold, triggering a special effect corresponding to the target gesture.

[0071] In an embodiment of the present application, when a user turns on the gesture recognition function, the system uses the gesture recognition model in the selected optimal target deep learning framework to perform gesture recognition. For example, it can recognize user gestures in real time at a preset frame rate of 15fps to obtain recognition results. In the recognition results, for example, if 3 out of 5 consecutive frames are recognized as the corresponding gesture, the recognition accuracy at this time is 60%, which is greater than the preset threshold, thereby triggering the corresponding gesture effect. For example, if 3 out of 5 consecutive video frames are recognized as "heart", the current page will trigger a heart-shaped special effect.

[0072] The following illustrates the overall process of the front-end gesture recognition system described above, using Figure 3. In Figure 3, when a user opens a webpage, the gesture recognition system initializes the user environment to determine whether the current environment supports Wasm. Depending on the determination, the system then proceeds to different subsequent processes. If the current operating environment does not support Wasm, the system automatically switches to a backend service, invoking the gesture recognition service deployed on the server via an HTTP interface. If the current operating environment supports Wasm, the system automatically performs idle state detection on the webpage, loading both the NCNN model and the Openvino model. These models are then tested for idle state, determining their resource consumption. The NCNN model's resource consumption is referred to as Resource Consumption 1, and the Openvino model's resource consumption is referred to as Resource Consumption 2. Resource Consumption 1 and Resource Consumption 2 are compared. If Resource Consumption 1 is greater than Resource Consumption 2, the Openvino model is retained and the NCNN model is uninstalled. If Resource Consumption 1 is less than Resource Consumption 2, the Openvino model is uninstalled and the NCNN model is retained. When the user turns on the gesture recognition function, gesture recognition is performed from the video stream at a preset frame rate of 15fps based on the final retained model, and the recognition result is obtained. When the user does not turn on the gesture recognition function, the current system is exited.

[0073] In the above-mentioned gesture recognition method, the gesture recognition model is obtained by: inputting training data into a first original model and then training it through a momentum encoder to obtain a first feature extraction model, wherein the training data is unlabeled image data; cross-entropy training is performed on a second original model based on the first feature extraction model and the training data to obtain a second feature extraction model, wherein the second original model is a simplified model of the first original model; and the second feature extraction model is trained using labeled data to obtain a gesture recognition model, wherein the gesture recognition model includes a hand detection model and a gesture classification model.

[0074] In the above steps, cross-entropy training is performed on the second original model based on the first feature extraction model and the training data to obtain the second feature extraction model, which specifically includes the following steps: inputting the training data into the first feature extraction model and the second original model respectively to obtain a first output and a second output; storing the first output and the historical output of the first feature extraction model into an instance queue; performing an inner product operation on the output in the instance queue and the first output to obtain a first probability distribution; performing an inner product operation on the output in the instance queue and the second output to obtain a second probability distribution; performing a cross-entropy operation on the first probability distribution and the second probability distribution, and training the second original model based on the result of the cross-entropy operation to obtain a second feature extraction model.

[0075] In an embodiment of the present application, unlabeled image data is used, and a contrastive learning method based on MoCoV2 is used to train a first feature extraction model with a backbone network of Resnet101. Specifically, positive samples are generated using data enhancement methods such as random cropping, and the others are negative samples. With the help of momentum encoder (i.e., the above-mentioned momentum encoder) and queue, end-to-end self-supervised training is completed. The self-supervised contrastive learning method has a good effect on large model training, but the effect on small model training is not ideal. The gesture recognition system has high requirements for model size and real-time performance. Therefore, a self-supervised distillation method is used, and the above-mentioned first feature extraction model is used as the teacher model, MobilenetV2-0.25 is used as the student model. An instance queue is constructed to store the output of the teacher network, the teacher network is frozen (i.e., the parameters of the teacher model are not updated), and the same unlabeled data is used. Cross-entropy training is performed by CrossEntropy loss, and the above-mentioned first feature extraction model (Resnet101) is distilled to obtain a small network (i.e., the second feature extraction model). This step can train the feature extraction model without labeling data, effectively reducing data labeling cost and time. The self-supervised distillation process is described below with reference to Figure 4.

[0076] In Figure 4, the unlabeled image data is input into the first feature extraction model (Resnet101) and the second original model (MobileNetV2), and the first output Z is obtained. T and the second output Z S , all outputs of the first feature extraction model are stored in the instance queue, including historical outputs and current outputs, the output in the instance queue is inner-producted with the first output, and a first probability distribution is obtained through a normalized exponential function, the output in the instance queue is inner-producted with the second output, and a second probability distribution is obtained through a normalized exponential function, the first probability distribution and the second probability distribution are cross-entropy operated, and the second original model is trained based on the result of the cross-entropy operation to obtain a second feature extraction model.

[0077] In the above steps, the second feature extraction model is trained using the labeled data to obtain a gesture recognition model, which specifically includes the following steps: using the second feature extraction model as the backbone network, freezing the feature extraction layer of the second feature extraction model, and adding a multi-layer perceptron layer after the feature extraction layer to obtain a detection network; training the detection network based on the first type of data in the labeled data to obtain a hand detection model, wherein the hand detection model is used to identify whether it is a gesture; training the detection network based on the second type of data in the labeled data to obtain a gesture classification model, wherein the gesture classification model is used to identify the category of the gesture.

[0078] In an embodiment of the present application, a hand detection model is trained based on the NanoDet method. Specifically, the second feature extraction model (MobilenetV2-0.25) obtained by distillation is used as the backbone network (Backbone), its feature extraction layer is frozen, only the gradient information of the feature extraction layer is back-transmitted, and the parameters are not updated. After the feature extraction layer, the head layer (conv+relu+bn) and the predict layer (ie, the MLP layer) are added to build a detection network. About 1,000 hand annotated data (ie, the first type of data mentioned above) are used to update the head layer and predict layer parameters to train a hand detection model.

[0079] Using cross-entropy as the loss, we used the distilled MobileNetV2-0.25 as the backbone network. We froze the feature extraction layer and only backpropagated the gradient information of the feature extraction layer without updating the parameters. We added an MLP layer after the feature extraction layer. We used a small amount of labeled data, approximately 1,000 gesture-labeled images (the second category of data mentioned above), to train the MLP layer and update the parameters of the head and predict layers to obtain a gesture classification model.

[0080] Through the above steps, a model equivalent to the traditional deep learning method can be trained using only thousands of labeled data, greatly reducing the labeling cost.

[0081] After obtaining the gesture recognition model in the above steps, the method further includes the following steps: converting the gesture recognition model into a gesture recognition model supported by the first deep learning framework based on a conversion tool of the first deep learning framework, thereby obtaining the gesture recognition model in the first deep learning framework; converting the gesture recognition model into a gesture recognition model supported by the second deep learning framework based on a conversion tool of the second deep learning framework, thereby obtaining the gesture recognition model in the second deep learning framework; encapsulating the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework into a software development kit using a target language; and converting the code in the software development kit into a target format using a format conversion tool.

[0082] In an embodiment of the present application, the conversion tools of NCNN and OpenVINO are used respectively to convert the hand detection model and the gesture classification model (collectively referred to as the gesture recognition model) into formats supported by the NCNN and OpenVINO frameworks, i.e., to obtain the ncnn model and the openvino model. Use C++ (i.e., the above-mentioned target language) to build a gesture recognition SDK, including color space conversion, cropping, and normalization of the input image frame, and send it to the hand detection model. After maximum normalization and other operations, the hand ROI area is obtained; then the hand ROI is sent to the gesture classification model to obtain the category corresponding to each gesture. Use the Emscripten tool (i.e., the above-mentioned format conversion tool) to convert the C++ code into wasm format (i.e., the above-mentioned target format) so that it can be deployed on the web page, transferring the computing pressure from the server segment to each user, and saving server costs.

[0083] The gesture recognition system is built using JavaScript. The JavaScript code implements video frame acquisition, gesture recognition WASM module calls, and the corresponding logic chain based on the gesture recognition results, displaying the corresponding gesture recognition effects. The following, combined with Figure 5, illustrates the process of building a front-end gesture recognition system based on self-supervised distillation.

[0084] In Figure 5, using unlabeled image data and a MoCoV2-based contrastive learning method, the first feature extraction model (or large model) with a Resnet101 backbone network is trained. The small model, MobilenetV2, is obtained through the self-supervised distillation process described above. The hand detection model and gesture classification model are trained using the NanoDet method and cross-entropy loss, respectively. Training the hand detection model and gesture classification model requires only a small amount of labeled data. The NCNN and OpenVINO conversion tools are used to convert the hand detection model and gesture classification model (collectively referred to as the gesture recognition model) into formats supported by the NCNN and OpenVINO frameworks, respectively, resulting in the NCNN and OpenVINO models. A gesture recognition SDK is built using C++ (the target language mentioned above). The C++ code is converted to the Wasm format using the Emscripten tool and deployed on the web.

[0085] The gesture recognition method provided by the embodiments of the present application has the following advantages: 1. Utilizing contrastive learning and self-supervised distillation methods, a feature extraction model is trained using unlabeled data, which is then distilled to a lightweight feature extraction model. This feature extraction model is then used as the backbone network for migration to gesture recognition tasks, reducing data annotation costs and development costs. 2. Utilizing WAS technology, deep learning models are deployed to web pages, significantly reducing server overhead and shifting intensive computing tasks to user devices, reducing machine costs. 3. Multiple options are provided, allowing users to adaptively select different models or services based on the user environment, selecting the optimal solution for each environment.

[0086] FIG6 is a structural diagram of a gesture recognition device according to an embodiment of the present application. As shown in FIG6 , the device includes:

[0087] The processing module 60 is configured to initialize the current operating environment in response to the start-up operation of the target object on the web page;

[0088] An operating module 62 is configured to respectively operate the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework according to the current operating environment, if the current operating environment supports the target format, to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the webpage;

[0089] a determination module 64 for determining a gesture recognition model in a target deep learning framework based on a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework;

[0090] The recognition module 66 is used to perform gesture recognition based on the gesture recognition model in the target deep learning framework.

[0091] By means of the processing module 60, the running module 62, the determining module 64, and the identifying module 66 in the above-mentioned gesture recognition device, the purpose of deploying the gesture recognition model on the web page is achieved, thereby achieving the technical effect of allocating the server cost to each user and saving the server overhead. This further solves the technical problem in related technologies of using gesture recognition solutions based on deep learning and deploying the model on the cloud, which consumes a large amount of server resources and costs.

[0092] In the processing module in the above-mentioned gesture recognition device, the processing module is also used to call the gesture recognition model deployed in the server in the form of a backend service when the current operating environment does not support the target format, wherein the gesture recognition model is used to recognize the gesture of the target object.

[0093] In the determination module in the above-mentioned gesture recognition device, the determination module is used to, when the first resource consumption is greater than the second resource consumption, use the gesture recognition model in the second deep learning framework as the gesture recognition model in the target deep learning framework and release the gesture recognition model in the first deep learning framework; when the first resource consumption is less than the second resource consumption, use the gesture recognition model in the first deep learning framework as the gesture recognition model in the target deep learning framework and release the gesture recognition model in the second deep learning framework; when the first resource consumption is equal to the second resource consumption, use the gesture recognition model in the first deep learning framework or the second deep learning framework as the gesture recognition model in the target deep learning framework and release the gesture recognition model outside the target deep learning framework.

[0094] In the recognition module in the above-mentioned gesture recognition device, the recognition module is used to recognize the gesture of the target object according to a preset frame rate to obtain a recognition result; when the accuracy of the target gesture recognized in the recognition result is greater than a preset threshold, the special effect corresponding to the target gesture is triggered.

[0095] The above-mentioned gesture recognition device also includes a training module 68, which is used to train a gesture recognition model. Specifically, the gesture recognition model is obtained by: inputting training data into a first original model and then training it through a momentum encoder to obtain a first feature extraction model, wherein the training data is unlabeled image data; cross-entropy training is performed on a second original model based on the first feature extraction model and the training data to obtain a second feature extraction model, wherein the second original model is a simplified model of the first original model; and the second feature extraction model is trained using labeled data to obtain a gesture recognition model, wherein the gesture recognition model includes a hand detection model and a gesture classification model.

[0096] In the training module in the above-mentioned gesture recognition device, the training module is also used to input training data into the first feature extraction model and the second original model respectively to obtain a first output and a second output; store the first output and the historical output of the first feature extraction model into an instance queue; perform an inner product operation on the output in the instance queue and the first output to obtain a first probability distribution; perform an inner product operation on the output in the instance queue and the second output to obtain a second probability distribution; perform a cross-entropy operation on the first probability distribution and the second probability distribution, and train the second original model based on the result of the cross-entropy operation to obtain a second feature extraction model.

[0097] In the training module of the above-mentioned gesture recognition device, the training module is also used to use the second feature extraction model as the backbone network, freeze the feature extraction layer of the second feature extraction model, and add a multi-layer perceptron layer after the feature extraction layer to obtain a detection network; the detection network is trained based on the first type of data in the label data to obtain a hand detection model, wherein the hand detection model is used to identify whether it is a gesture; the detection network is trained based on the second type of data in the label data to obtain a gesture classification model, wherein the gesture classification model is used to identify the category of the gesture.

[0098] In the processing module in the above-mentioned gesture recognition device, the processing module is also used to convert the gesture recognition model into a gesture recognition model supported by the first deep learning framework based on the conversion tool of the first deep learning framework, thereby obtaining the gesture recognition model in the first deep learning framework; convert the gesture recognition model into a gesture recognition model supported by the second deep learning framework based on the conversion tool of the second deep learning framework, thereby obtaining the gesture recognition model in the second deep learning framework; encapsulate the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework into a software development kit using the target language; and convert the code in the software development kit into the target format using the format conversion tool.

[0099] It should be noted that the gesture recognition device shown in FIG6 is used to execute the gesture recognition method shown in FIG2 , so the relevant explanations in the above gesture recognition method are also applicable to the gesture recognition device and will not be repeated here.

[0100] An embodiment of the present application also provides an electronic device, including: a memory for storing program instructions; a processor, connected to the memory, for executing program instructions to implement the following functions: initializing the current operating environment in response to a startup operation of a target object on a web page; when the current operating environment supports a target format, respectively running a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operating environment to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; determining the gesture recognition model in the target deep learning framework based on a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; and performing gesture recognition according to the gesture recognition model in the target deep learning framework.

[0101] It should be noted that the above electronic device is used to execute the gesture recognition method shown in FIG. 2 , so the relevant explanations in the above gesture recognition method are also applicable to the electronic device and will not be repeated here.

[0102] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located performs the following gesture recognition method by running the computer program: in response to the startup operation of the target object on the web page, initializing the current operating environment; when the current operating environment supports the target format, running the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework respectively according to the current operating environment to obtain the first resource consumption and the second resource consumption, wherein the target format is the format supported by the web page; based on the comparison result of the first resource consumption and the second resource consumption, determining the gesture recognition model in the target deep learning framework, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; performing gesture recognition according to the gesture recognition model in the target deep learning framework.

[0103] It should be noted that the above-mentioned non-volatile storage medium is used to execute the gesture recognition method shown in FIG. 2 , so the relevant explanations in the above-mentioned gesture recognition method are also applicable to the non-volatile storage medium, which will not be repeated here.

[0104] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0105] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0106] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0107] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0108] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0109] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0110] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for gesture recognition, comprising: In response to a start-up operation of the target object on the web page, initializing the current operating environment; In a case where the current operating environment supports the target format, respectively running a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operating environment to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; Determining a gesture recognition model in a target deep learning framework according to a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; Perform gesture recognition based on the gesture recognition model in the target deep learning framework.

2. The method according to claim 1, further comprising: After initializing the current running environment, In the case that the current operating environment does not support the target format, a gesture recognition model deployed in the server is called in the form of a backend service, wherein the gesture recognition model is used to recognize the gesture of the target object.

3. The method according to claim 1, wherein: Determining a gesture recognition model in a target deep learning framework according to a comparison result of the first resource consumption and the second resource consumption includes: When the first resource consumption is greater than the second resource consumption, using the gesture recognition model in the second deep learning framework as the gesture recognition model in the target deep learning framework, and releasing the gesture recognition model in the first deep learning framework; When the first resource consumption is less than the second resource consumption, using the gesture recognition model in the first deep learning framework as the gesture recognition model in the target deep learning framework, and releasing the gesture recognition model in the second deep learning framework; When the first resource consumption is equal to the second resource consumption, the gesture recognition model in the first deep learning framework or the second deep learning framework is used as the gesture recognition model in the target deep learning framework, and the gesture recognition model outside the target deep learning framework is released.

4. The method according to claim 1, wherein: Performing gesture recognition according to the gesture recognition model in the target deep learning framework includes: Recognize the gesture of the target object according to a preset frame rate to obtain a recognition result; When the accuracy of the target gesture identified in the recognition result is greater than a preset threshold, a special effect corresponding to the target gesture is triggered.

5. The method according to claim 1, wherein: The gesture recognition model is obtained in the following way: The training data is input into the first original model and then trained through the momentum encoder to obtain the first feature extraction Model, wherein the training data is unlabeled image data; Performing cross entropy training on a second original model according to the first feature extraction model and the training data to obtain a second feature extraction model, wherein the second original model is a simplified model of the first original model; The second feature extraction model is trained using the label data to obtain the gesture recognition model, wherein the gesture recognition model includes a hand detection model and a gesture classification model.

6. The method according to claim 5, wherein: Performing cross entropy training on the second original model according to the first feature extraction model and the training data to obtain a second feature extraction model, including: Inputting the training data into the first feature extraction model and the second original model respectively to obtain a first output and a second output; storing the first output and the historical output of the first feature extraction model in an instance queue; Perform an inner product operation on the output in the instance queue and the first output to obtain a first probability distribution; Perform an inner product operation on the output in the instance queue and the second output to obtain a second probability distribution; A cross entropy operation is performed on the first probability distribution and the second probability distribution, and the second original model is trained according to the result of the cross entropy operation to obtain the second feature extraction model.

7. The method according to claim 5, wherein: The second feature extraction model is trained using the label data to obtain the gesture recognition model, including: The second feature extraction model is used as a backbone network, and a feature extraction layer of the second feature extraction model is frozen, and a multi-layer perceptron layer is added after the feature extraction layer to obtain a detection network; Training the detection network according to the first type of data in the label data to obtain the hand detection model, wherein the hand detection model is used to identify whether it is a gesture; The detection network is trained according to the second type of data in the label data to obtain the gesture classification model, wherein the gesture classification model is used to identify the category of the gesture.

8. The method according to claim 5, further comprising: After obtaining the gesture recognition model, According to a conversion tool of the first deep learning framework, converting the gesture recognition model into a gesture recognition model supported by the first deep learning framework to obtain a gesture recognition model in the first deep learning framework; According to a conversion tool of the second deep learning framework, converting the gesture recognition model into a gesture recognition model supported by the second deep learning framework, to obtain a gesture recognition model in the second deep learning framework; Encapsulating the gesture recognition model in the first deep learning framework and the gesture recognition model in the second deep learning framework into a software development kit using a target language; A format conversion tool is used to convert the code in the software development kit into the target format.

9. The method according to claim 1, wherein: In response to the start-up operation of the target object on the web page, the current operating environment is initialized, including: Determine whether the current operating environment supports the target format; When the current operating environment does not support the target format, the backend service is switched to call the gesture recognition service deployed in the server through the http interface; When the current operating environment supports the target format, idle detection is performed on the web page.

10. A gesture recognition device, comprising: A processing module, used to initialize the current operating environment in response to a start-up operation of the target object on the web page; an operation module, configured to respectively operate a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operation environment when the current operation environment supports a target format, to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; A determination module, configured to determine a gesture recognition model in a target deep learning framework according to a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; A recognition module is used to perform gesture recognition based on a gesture recognition model in the target deep learning framework.

11. An electronic device, comprising: A memory for storing program instructions; A processor is connected to the memory and is used to execute program instructions that implement the following functions: in response to a startup operation of a target object on a web page, initializing a current operating environment; when the current operating environment supports a target format, respectively running a gesture recognition model in a first deep learning framework and a gesture recognition model in a second deep learning framework according to the current operating environment to obtain a first resource consumption and a second resource consumption, wherein the target format is a format supported by the web page; determining a gesture recognition model in a target deep learning framework according to a comparison result of the first resource consumption and the second resource consumption, wherein the target deep learning framework is the first deep learning framework or the second deep learning framework; and performing gesture recognition according to the gesture recognition model in the target deep learning framework.

12. A non-volatile storage medium comprising a stored computer program, wherein: The device where the non-volatile storage medium is located executes the method for gesture recognition described in any one of claims 1 to 9 by running the computer program.

Citation Information

Patent Citations

  • Gesture recognition method and system based on knowledge distillation and attention mechanism

    CN113449610A

  • Terminal interface test method and device

    CN115269359A

  • Human hand posture recognition system and method based on TinyML technology

    CN115984903A

  • Gesture recognition method and device and electronic equipment

    CN117648033A

  • Smart device

    US20170312614A1