Gesture recognition method, apparatus, device, and storage medium

By capturing specified gestures to activate the gesture recognition function, and combining low-power and high-power image acquisition modes, accurate processing and feedback of gesture information are achieved. This solves the problem of resource waste in device wake-up and interaction, provides sign language translation and command conversion functions, meets the communication needs of deaf and mute people, and improves the naturalness and efficiency of device interaction.

CN122111209APending Publication Date: 2026-05-29BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-11-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing gesture recognition technology suffers from resource waste and false or missed triggers during device wake-up and interaction, and lacks effective sign language recognition and translation functions, making it difficult to meet the communication needs of deaf and mute individuals.

Method used

The gesture recognition function is activated by capturing a specified gesture. It uses a low-power image acquisition mode to monitor the appearance of the gesture, switches to a high-power mode for accurate recognition, and combines sign language translation, command conversion and voice conversion functions to achieve accurate processing and feedback of gesture information.

Benefits of technology

It improves the accuracy and efficiency of gesture recognition, provides a non-verbal communication method, meets the interactive needs of deaf and mute people, enhances the natural fluency and personalized operation of device interaction, and promotes the barrier-free flow of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111209A_ABST
    Figure CN122111209A_ABST
Patent Text Reader

Abstract

The present disclosure provides a gesture recognition method and device, equipment and storage medium, and relates to the technical field of computer. The method comprises the following steps: in response to capturing a specified gesture, waking up a gesture recognition function corresponding to the specified gesture; in response to receiving gesture information, processing the gesture information through the gesture recognition function, and displaying the processing result. The method can quickly wake up and operate the device through a simple gesture, without complicated key or touch operation, making the interaction more natural and smooth. Through the process of waking up first and then recognizing, the processing of invalid gestures can be effectively reduced, the accuracy and efficiency of gesture recognition can be improved, and it is ensured that the device can accurately respond to the user's instruction. In addition, it also provides a non-verbal communication method for hearing-impaired or speech-impaired people, helping them to interact more smoothly with intelligent devices, and promoting the barrier-free circulation of information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a gesture recognition method, apparatus, device, and storage medium. Background Technology

[0002] Gesture recognition technology, especially for sign language recognition, is increasingly becoming an important development direction in the field of human-computer interaction. Sign language is not only an important means of communication for deaf and mute people, but it also demonstrates value in many fields such as education and entertainment. Through accurate gesture and sign language recognition, it is possible not only to promote barrier-free communication between deaf and mute people and hearing people, but also to create richer interactive experiences in fields such as virtual reality and gaming.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this disclosure is to provide a gesture recognition method, apparatus, device, and storage medium.

[0005] According to a first aspect of the present disclosure, a gesture recognition method is provided, comprising: in response to capturing a specified gesture, activating a gesture recognition function corresponding to the specified gesture; in response to receiving gesture information, processing the gesture information through the gesture recognition function, and displaying the processing result.

[0006] In some embodiments, the gesture recognition method further includes: capturing the specified gesture through a first image acquisition mode; and receiving the gesture information through a second image acquisition mode in response to the gesture recognition function being activated; wherein the power consumption of the first image acquisition mode is lower than the power consumption of the second image acquisition mode.

[0007] In some embodiments, the gesture recognition function includes a sign language translation function; the gesture information is sign language information; wherein, in response to receiving the gesture information, the gesture recognition function processes the gesture information and displays the processing result, including: in response to receiving the sign language information, the sign language translation function translates the sign language information to obtain translated text information; and displays and / or broadcasts the translated text information.

[0008] In some implementations, translating the sign language information using the sign language translation function to obtain translated text information includes: translating the sign language information using the sign language translation function to obtain initial text information; generating multiple fuzzy recommended text information based on the initial text information; displaying the multiple fuzzy recommended text information; receiving a selection operation on the multiple fuzzy recommended text information; and determining the translated text information based on the selected fuzzy recommended text information.

[0009] In some embodiments, the gesture recognition function includes an instruction conversion function; wherein, in response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result, including: in response to receiving gesture information, converting the gesture information into a gesture instruction through the instruction conversion function; controlling the object of the gesture instruction to execute the gesture instruction; and displaying the execution result of the gesture instruction.

[0010] In some embodiments, the gesture information is sign language information; wherein, in response to receiving gesture information, converting the gesture information into a gesture command through the command conversion function includes: in response to receiving sign language information, parsing the sign language information through the command conversion function to determine the target and processing method indicated by the sign language information; and generating the gesture command according to the target and processing method.

[0011] In some implementations, in response to receiving gesture information, the gesture information is converted into a gesture command through the command conversion function, including: in response to receiving gesture information, calling the command library associated with the command conversion function; and in response to the gesture information matching a target preset command in the command library, determining the preset command as the gesture command.

[0012] In some embodiments, the gesture recognition function includes a speech conversion function; the gesture information is sign language information; wherein, in response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result, including: in response to receiving sign language information, processing the sign language information through the speech conversion function to obtain semantic information corresponding to the sign language information; generating audio information based on the semantic information; and playing the audio information.

[0013] According to a second aspect of the present disclosure, a gesture recognition device is provided, comprising: a wake-up unit, configured to wake up a gesture recognition function corresponding to the specified gesture in response to capturing a specified gesture; and a processing unit, configured to process the gesture information through the gesture recognition function in response to receiving gesture information, and display the processing result.

[0014] In some embodiments, the wake-up unit is further configured to: capture the specified gesture through a first image acquisition mode; the processing unit is further configured to: receive the gesture information through a second image acquisition mode in response to the gesture recognition function being woken up; wherein the power consumption of the first image acquisition mode is lower than the power consumption of the second image acquisition mode.

[0015] In some embodiments, the gesture recognition function includes a sign language translation function; the gesture information is sign language information; wherein, in response to receiving the gesture information, the processing unit processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving the sign language information, translating the sign language information through the sign language translation function to obtain translated text information; and displaying and / or broadcasting the translated text information.

[0016] In some embodiments, the processing unit is further configured to: translate the sign language information using the sign language translation function to obtain initial text information; generate multiple fuzzy recommended text information based on the initial text information; display the multiple fuzzy recommended text information; receive a selection operation on the multiple fuzzy recommended text information; and determine the translated text information based on the selected fuzzy recommended text information.

[0017] In some embodiments, the gesture recognition function includes an instruction conversion function; wherein, in response to receiving gesture information, the processing unit processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving gesture information, converting the gesture information into a gesture instruction through the instruction conversion function; controlling the object of the gesture instruction to execute the gesture instruction; and displaying the execution result of the gesture instruction.

[0018] In some embodiments, the gesture information is sign language information; wherein, in response to receiving the gesture information, the processing unit converts the gesture information into a gesture command through the command conversion function, including: in response to receiving the sign language information, parsing the sign language information through the command conversion function to determine the target and processing method indicated by the sign language information; and generating the gesture command according to the target and processing method.

[0019] In some implementations, the processing unit, in response to receiving gesture information, converts the gesture information into a gesture command through the command conversion function, including: in response to receiving gesture information, calling the command library associated with the command conversion function; and in response to the gesture information matching a target preset command in the command library, determining the preset command as the gesture command.

[0020] In some embodiments, the gesture recognition function includes a speech conversion function; the gesture information is sign language information; wherein, in response to receiving the gesture information, the processing unit processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving the sign language information, processing the sign language information through the speech conversion function to obtain semantic information corresponding to the sign language information; generating audio information based on the semantic information; and playing the audio information.

[0021] According to a third aspect of the present disclosure, an electronic device is provided, characterized in that it includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the gesture recognition method described above.

[0022] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute a gesture recognition method, the method comprising: activating a gesture recognition function corresponding to the specified gesture in response to capturing a specified gesture; and processing the gesture information through the gesture recognition function in response to receiving gesture information, and displaying the processing result.

[0023] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the gesture recognition method described above.

[0024] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0025] This disclosure allows for quick activation and operation of devices with simple gestures, eliminating the need for cumbersome button or touch operations, resulting in a more natural and fluid interaction. The wake-up-then-recognition process effectively reduces the processing of invalid gestures, improving the accuracy and efficiency of gesture recognition and ensuring the device responds accurately to user commands. Furthermore, it provides a non-verbal communication method for hearing-impaired or speech-impaired individuals, helping them interact more smoothly with smart devices, promoting barrier-free information flow, and demonstrating broad application prospects and social value.

[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0027] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0028] Figure 1This is a flowchart illustrating a gesture recognition method according to some embodiments of the present disclosure.

[0029] Figure 2 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0030] Figure 3 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0031] Figure 4 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0032] Figure 5 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0033] Figure 6 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0034] Figure 7 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0035] Figure 8 This is a schematic diagram of a system that can implement a gesture recognition method according to some embodiments of the present disclosure.

[0036] Figure 9 This is a block diagram illustrating a gesture recognition device according to some embodiments of the present disclosure.

[0037] Figure 10 This is a block diagram illustrating a device for gesture recognition according to some embodiments of the present disclosure. Detailed Implementation

[0038] Exemplary embodiments of this disclosure will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.

[0039] The embodiments described below, which are examples of some of the embodiments of this disclosure, do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0040] The specific implementation methods of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0041] Figure 1 This is a flowchart illustrating a gesture recognition method according to some embodiments of the present disclosure, such as... Figure 1 As shown, the gesture recognition method can be applied to terminals with image acquisition and image processing capabilities. The terminal may include, but is not limited to, smartphones, tablets, laptops, desktop computers, augmented reality devices, virtual reality devices, etc. The gesture recognition method may include the following steps.

[0042] In step S110, in response to capturing a specified gesture, the gesture recognition function corresponding to the specified gesture is activated.

[0043] In this embodiment, a listening mechanism can be set to capture user gestures. When a specific gesture (i.e., a designated gesture) is captured, the gesture recognition module or function associated with that gesture can be activated. Since the user performs the designated gesture, it means that the user has a need to use the relevant gesture recognition module. Therefore, this embodiment ensures that the corresponding gesture recognition is only activated when the user needs it, saving resources and achieving accurate function activation, avoiding false triggering or missed triggering.

[0044] In this embodiment of the disclosure, a camera device (such as a webcam), an infrared sensor, radar, or other types of sensors can be used to capture the user's hand movements. The camera device can directly capture visual information, facilitating subsequent processing; the camera device can include, but is not limited to, ordinary cameras, event cameras, TOF cameras, and other camera devices capable of capturing motion and depth information of gestures.

[0045] In this embodiment of the disclosure, a corresponding designated gesture can be predefined for each gesture recognition module. These gestures can be clear, easily recognizable, and user-friendly. For example, waving, clenching a fist, making a heart shape with hands, drawing a circle, and opening the palm.

[0046] In this embodiment of the disclosure, when the gesture recognition function is activated, the corresponding gesture recognition function module, such as image processing algorithm, machine learning model, etc., can be called; the gesture recognition function module can be located on the same device as the gesture capture function module, or it can be located on a different device.

[0047] In an exemplary embodiment, both the gesture capture and gesture recognition modules can be integrated into a smartphone or tablet. This integration simplifies the overall system architecture and reduces data transmission latency and complexity.

[0048] In an exemplary embodiment, the gesture capture module can be located on one device (such as a wearable device, camera, etc.), while the gesture recognition module can be located on another device (such as a server, cloud, computing center, etc.). This distributed approach can fully utilize the computing resources and storage capabilities of different devices to achieve more efficient and flexible gesture recognition.

[0049] In an exemplary embodiment, when a user makes a specified gesture corresponding to a different gesture recognition function, multiple gesture recognition functions can be activated simultaneously.

[0050] In an exemplary embodiment, a clear feedback mechanism can also be provided to let users know whether their gestures have been correctly recognized. For example, after a specified gesture is captured, the gesture recognition function corresponding to the specified gesture can be determined first, and a prompt message indicating that the function will be activated can be displayed. When the user determines that the function is the one they need to activate, the gesture recognition function can be activated.

[0051] In some embodiments of this disclosure, the gesture recognition method may further include: disabling the gesture recognition function in response to capturing a function-off gesture.

[0052] The function-off gesture can be used for different types of gesture recognition functions, or different function-off gestures can be pre-set for each gesture recognition function.

[0053] In step S120, in response to receiving gesture information, the gesture information is processed through the gesture recognition function, and the processing result is displayed.

[0054] In this embodiment of the disclosure, after the gesture recognition function is activated, it can begin to receive data containing gesture information, and process the gesture information through the gesture recognition function to determine the specific meaning or type of the gesture. The gesture information may include the trajectory, type, position, direction, and duration of the gesture.

[0055] In this embodiment of the disclosure, based on the result of gesture recognition, corresponding information can be displayed or corresponding operations can be performed on the user interface. Specifically, after processing gesture information through the gesture recognition function, the processing result can be understood semantic information, generated instructions, etc. The display of the processing result can be, for example, displayed on the screen, through voice feedback, haptic feedback, or as the result of instruction execution. This feedback mechanism of displaying processing results can enhance the user's interactive experience, making it easier for the user to understand and confirm the system's response.

[0056] For example, if a user makes a "clenched fist" gesture, the gesture recognition function can identify the gesture and determine that it corresponds to a command to open a specific application, then the command can be executed directly. As another example, if the recognized gesture is "turn on the air conditioner," the system may automatically turn on the associated air conditioner.

[0057] In an exemplary embodiment, the gesture information may be sign language gestures, and the gesture recognition function can be implemented by a sign language detection algorithm to identify the semantics corresponding to the sign language gestures made by the user; wherein, the sign language detection algorithm may be implemented based on a convolutional neural network (CNN) or a recurrent neural network (RNN).

[0058] This embodiment of the invention does not rely on specific hardware devices or complex settings, and has high versatility and compatibility. It can be widely applied in multiple fields such as smartphones, smart homes, and wearable devices, providing a convenient gesture interaction method for various smart devices.

[0059] As can be seen from the above steps, the gesture recognition method provided in this disclosure can first activate the corresponding gesture recognition function based on the captured specified gesture, then process the gesture information received after the function is activated, and display the processing result. Therefore, the gesture recognition method provided in this disclosure, on the one hand, allows for quick activation and operation of the device with just a simple gesture, eliminating the need for cumbersome button or touch operations, making the interaction more natural and fluid; furthermore, it ensures that the gesture recognition function corresponding to the gesture is activated, saving resources and achieving accurate function activation, avoiding accidental or missed triggers; in addition, it allows users to use different gestures to perform different functions in different scenarios, thereby increasing the diversity and flexibility of interaction and meeting the personalized operation needs of users in different situations. On the other hand, the process of activating before recognizing effectively reduces the processing of invalid gestures, improves the accuracy and efficiency of gesture recognition, and ensures that the device can accurately respond to user commands. Moreover, for people with hearing or speech impairments, the gesture recognition method provided in this disclosure can provide a non-verbal communication method, helping them to interact more smoothly with smart devices, promoting the barrier-free flow of information, and demonstrating broad application prospects and social value.

[0060] In some embodiments of this disclosure, the gesture recognition method further includes: capturing the specified gesture through a first image acquisition mode; and receiving the gesture information through a second image acquisition mode in response to the gesture recognition function being activated; wherein the power consumption of the first image acquisition mode is lower than the power consumption of the second image acquisition mode.

[0061] In this embodiment of the disclosure, during the initial stage, i.e., when the gesture recognition function is not activated, a low-power first image acquisition mode can be used to continuously monitor the occurrence of gestures and capture the specified gesture after its occurrence is detected. In this first image acquisition mode, the image resolution, frame rate, or processing complexity can be appropriately reduced to decrease the demand for computing resources and power. This design helps maintain low-power operation of the device during long-term standby or monitoring states, extending battery life and reducing unnecessary energy consumption, especially when the device is in standby for extended periods or when the user does not frequently use gesture recognition.

[0062] In an exemplary embodiment, an event camera can be used to implement the first image acquisition mode, wherein the event camera has lower power consumption compared to a regular camera. When the event camera detects an event signal, it triggers event-based gesture detection, captures event frame information, and then determines the gesture recognition function to be activated based on the event frame information.

[0063] In an exemplary embodiment, the first image acquisition mode can be implemented using the low-power detection mode of a conventional camera. When the conventional camera detects a change in light, it can be triggered to capture one or more images in low-power mode, and then the gesture recognition function to be activated can be determined through the images.

[0064] In an exemplary embodiment, an event camera and a regular camera can also be integrated on the same device, and the usage priority of the low-power detection modes of the event camera and the regular camera can be determined. The usage priority can be set to the event camera or the regular camera required in different environments.

[0065] In this embodiment of the disclosure, after the gesture recognition function is activated, it can switch to a second image acquisition mode to receive more detailed gesture information. The second image acquisition mode has higher power consumption than the first image acquisition mode, but it also provides higher image quality and more refined gesture recognition capabilities. In the second image acquisition mode, the image resolution can be increased, the frame rate can be improved, or more complex image processing algorithms can be used to ensure accurate and rapid reception and processing of subsequent gesture information.

[0066] The dual-mode switching strategy provided in this disclosure ensures both low power consumption in standby mode and high-quality gesture recognition services when needed, thus achieving an optimized balance between power consumption and performance. Furthermore, it enhances the user experience, as users do not need to worry about the device rapidly depleting its battery during prolonged operation, while still receiving fast and accurate gesture responses when required.

[0067] Figure 2This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure. For example... Figure 2 As shown, the gesture recognition method may include the following steps.

[0068] Step S210: In response to capturing a specified gesture through the first image acquisition mode, the gesture recognition function corresponding to the specified gesture is activated.

[0069] Step S220: In response to receiving gesture information through the second image acquisition mode, the gesture information is processed through the gesture recognition function, and the processing result is displayed; wherein, the power consumption of the first image acquisition mode is lower than the power consumption of the second image acquisition mode.

[0070] The specific methods described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0071] Figure 3 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure. For example... Figure 3 As shown, the gesture recognition method may include the following steps.

[0072] Step S310: In response to capturing a specified gesture, the gesture recognition function corresponding to the specified gesture is activated; wherein, the gesture recognition function includes a sign language translation function.

[0073] In this embodiment of the disclosure, the sign language translation function can be used as a type of gesture recognition function, specifically for recognizing and translating sign language.

[0074] Step S320: In response to receiving sign language information, the sign language information is translated using the sign language translation function to obtain translated text information.

[0075] In this embodiment of the disclosure, when a user makes sign language gestures, this gesture information can be captured. Subsequently, a dedicated sign language translation function can be used to process this information. The function converts sign language into plain text information; that is, it translates the content expressed by the user through sign language into a written form that can be understood and read by most people. Specifically, the sign language translation function can utilize advanced algorithms and / or machine learning models to recognize and translate sign language.

[0076] Step S330: Display and / or broadcast the translated text information.

[0077] In this embodiment of the disclosure, once the sign language translation function has completed the translation, the translated text information can be displayed (including showing and / or broadcasting). In this way, even people who do not understand sign language can understand the content expressed by the user through sign language by reading the translated text.

[0078] In an exemplary embodiment, the translated text information may be displayed, broadcast, or simultaneously displayed and broadcast on a designated device. This device may be a device that implements sign language translation functionality, or it may be a different device from the device that implements sign language translation functionality.

[0079] For example, when the translation function is integrated into a device with a display (such as a smartphone, tablet, or a dedicated translator with a camera), the translated text can be displayed directly on the device's screen; this integrated display method simplifies the operation process, and users can view the translation results without additional equipment or steps.

[0080] For example, devices that enable sign language translation can send the translation results to other designated devices via wireless or wired connections. This separate display method increases the flexibility of use and can meet the needs of users who want to display translated text information on other designated devices (such as large screen displays, TVs, or projectors). It can be applied to various scenarios such as meetings, teaching, and speeches.

[0081] In an exemplary embodiment, after displaying the translated text information, a playback control may also be displayed, and then, in response to the user's triggering of the playback control, the voice corresponding to the translated text information is played.

[0082] Through the embodiments of this disclosure, the sign language translation function provides a convenient way for sign language users and others to communicate, making the transmission of information smoother and more efficient.

[0083] In some embodiments of this disclosure, the sign language information is translated using the sign language translation function to obtain translated text information, including: translating the sign language information using the sign language translation function to obtain initial text information; generating multiple fuzzy recommended text information based on the initial text information; displaying the multiple fuzzy recommended text information; receiving a selection operation on the multiple fuzzy recommended text information; and determining the translated text information based on the selected fuzzy recommended text information.

[0084] In this embodiment of the disclosure, the fuzzy recommendation text information may include synonyms, near-synonyms, related phrases, or possible alternative expressions, aiming to provide users with more diverse choices. For example, possible alternative or supplementary text may be generated based on keywords or context of the initial text information as fuzzy recommendation text information, or fuzzy recommendation text information may be generated based on semantic similarity and / or user's historical choices.

[0085] These fuzzy recommendation texts are displayed to the user so they can view and select the text that best matches their intent. The user can then select their preferred fuzzy recommendation text using interactive elements on the interface (such as buttons, touch areas, etc.). After receiving the user's selection, the system can determine the selected fuzzy recommendation text as the translation text.

[0086] In an exemplary embodiment, when sign language is detected, the sign language translation function can begin sign language comparison; and in order to simplify the difficulty of sign language input and improve the accuracy of input, fuzzy recommendations can be made after sign language input. For example, when the input sign language means "open", multiple fuzzy recommendation text information such as "1, open, 2, expand, 3, open, 4..." can be made. Then, the user can select by gesturing left or right, thereby accelerating sign language input.

[0087] By integrating sign language translation functionality with the generation and display of fuzzy recommended text information through the embodiments of this disclosure, not only can efficient sign language recognition and translation be achieved, but users are also provided with more diverse choices and a personalized communication experience, thereby improving the accuracy and efficiency of communication and enhancing the overall user experience.

[0088] In an exemplary embodiment, the initial text information may also be directly displayed and / or broadcast as translated text information.

[0089] Figure 4 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure. For example... Figure 4 As shown, the gesture recognition method may include the following steps.

[0090] Step S410: In response to capturing a specified gesture, the gesture recognition function corresponding to the specified gesture is activated; wherein, the gesture recognition function includes an instruction conversion function.

[0091] Step S420: In response to receiving gesture information, the gesture information is converted into a gesture command through the command conversion function.

[0092] In this embodiment of the disclosure, the instruction conversion function can convert gesture information into specific gesture instructions. These instructions can be commands to control the device, operations to select menu items, input text information, etc. This conversion may be based on a predefined rule base, machine learning model, or deep learning algorithm to map gesture information to specific instructions, ensuring the accuracy and efficiency of the conversion.

[0093] Step S430: Control the target of the gesture instruction to execute the gesture instruction.

[0094] In this embodiment of the disclosure, after the gesture information is converted into a gesture command, the system can find and control the corresponding target (such as a device, application, or interface element) according to the content of the command and the control logic. The control process may involve various operations such as calling the device's API, sending control signals, and modifying the application state.

[0095] In this embodiment of the disclosure, when the target is a device, the device can be the same as or different from the device that implements the instruction conversion function. When the target is a different device from the device that implements the instruction conversion function, a control signal can be sent to the target to control it.

[0096] Step S440: Display the execution result of the gesture command.

[0097] In this embodiment of the disclosure, the execution result of the gesture command can be displayed in an intuitive way, such as by updating the interface, playing a sound, or providing vibration feedback. In this way, the user can use this feedback to confirm whether the command has been executed and whether the result meets expectations.

[0098] Through the embodiments disclosed herein, an instruction conversion function can be introduced to convert gesture information into specific instructions and control the corresponding objects to execute these instructions. This not only improves the practicality and interactivity of gesture recognition, but also provides a more convenient and efficient solution for fields such as smart homes, gaming entertainment, and assisted communication.

[0099] Figure 5 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0100] In this embodiment of the disclosure, Figure 5 In the gesture recognition method shown, steps S510, S540, and S550 are respectively related to... Figure 4 Steps S410, S430, and S440 in the gesture recognition method shown correspond to each other and will not be repeated here.

[0101] In this embodiment of the disclosure, Figure 4 Based on the gesture recognition method shown, Figure 5 The gesture recognition method shown may also include the following steps.

[0102] Step S520: In response to receiving sign language information, the sign language information is parsed through the instruction conversion function to determine the target and processing method indicated by the sign language information.

[0103] In this embodiment, a user can input sign language (a specific gesture language) information into the second image acquisition mode. This sign language information can be passed to the instruction conversion function. The instruction conversion function can use a pre-trained model or algorithm to parse the sign language information and identify its specific meaning or instruction. Specifically, semantic understanding, context analysis, and other processing techniques can be used to determine the target (such as a smart home device, application, or interface element) and processing method (such as turning on, turning off, or adjusting) indicated in the sign language information.

[0104] Step S530: Generate the gesture command according to the target and processing method.

[0105] In this embodiment of the disclosure, specific instructions can be generated based on the parsed target and processing method. These instructions may be data encoded in a specific format, used to control devices or applications to perform corresponding operations.

[0106] For example, deaf and mute individuals can interact with smart devices using sign language, such as sending messages, making phone calls, and controlling home appliances. Furthermore, in special education, teachers can use sign language to interact with students, assisting them in learning and acquiring knowledge.

[0107] This disclosure introduces sign language information parsing and command conversion functions, enabling not only accurate recognition of sign language information but also its conversion into specific gesture commands, allowing corresponding objects to execute these commands. This not only enhances the practicality and interactivity of gesture recognition but also provides a more convenient and efficient communication method for people with language impairments, further promoting the development of accessibility technology.

[0108] Figure 6 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure.

[0109] In this embodiment of the disclosure, Figure 6 In the gesture recognition method shown, steps S610, S640, and S650 are respectively related to... Figure 4 Steps S410, S430, and S440 in the gesture recognition method shown correspond to each other and will not be repeated here.

[0110] In this embodiment of the disclosure, Figure 4 Based on the gesture recognition method shown, Figure 6 The gesture recognition method shown may also include the following steps.

[0111] Step S620: In response to receiving gesture information, the instruction library associated with the instruction conversion function is invoked.

[0112] In this embodiment, a command library can be pre-set, which can store a series of predefined gesture commands and their corresponding gesture information. When gesture information is received, this command library is maliciously invoked to perform subsequent matching operations.

[0113] Step S630: In response to the gesture information matching the target preset instruction in the instruction library, the preset instruction is determined as the gesture instruction.

[0114] In this embodiment, the received gesture information can be compared one by one with preset commands in the command library to find a match. The matching process can be based on various technologies such as image recognition, feature extraction, and template matching to ensure accuracy and efficiency. Once a match is found, the corresponding preset command can be identified as the current gesture command. This gesture command will then be used to control the behavior of the device or application to achieve the user's intent.

[0115] In an exemplary embodiment, if there are multiple matching results, multiple preset matching instructions can be displayed, and then the final gesture instruction is determined based on the user's selection. For example, if the gesture information corresponds to the meaning of "turn on the light", when there are multiple controllable lights, multiple instruction options such as "1. Turn on the living room light, 2. Turn on the dining room light, 3. Turn on the bedroom light, 4..." can be displayed to the user. The user can then select the option by gesturing left or right or by touching, thereby determining the final gesture instruction.

[0116] For example, users can control smart home devices with simple gestures, such as turning on lights or adjusting the air conditioning temperature; the system can recognize and execute these gesture commands based on a preset command library. As another example, while driving, users can control various functions of the in-vehicle system with gestures, such as changing songs or adjusting the volume; the preset command library ensures that these gesture commands are accurately recognized and executed.

[0117] Through the embodiments disclosed herein, gesture commands can be generated by calling a command library and matching preset commands, providing users with a more intuitive and convenient operation method. This design is not only suitable for smart homes and gaming entertainment, but can also be widely applied in various fields such as barrier-free communication, educational assistance, and medical rehabilitation, providing more intelligent and user-friendly services for users with different needs.

[0118] Figure 7 This is a flowchart illustrating yet another gesture recognition method according to some embodiments of the present disclosure. For example... Figure 7 As shown, the gesture recognition method may include the following steps.

[0119] Step S710: In response to capturing a specified gesture, the gesture recognition function corresponding to the specified gesture is activated; wherein, the gesture recognition function includes a speech conversion function.

[0120] Step S720: In response to receiving sign language information, the sign language information is processed through the speech conversion function to obtain semantic information corresponding to the sign language information.

[0121] In this embodiment of the disclosure, after confirming that the received information is sign language, a speech-to-text function can be used to process this sign language information. The user's sign language gestures are a silent expression of the user's intentions.

[0122] In an exemplary practical example, the integrated speech conversion function in the system can incorporate technologies such as image recognition, motion capture, and natural language processing to parse the received sign language information. The speech conversion function can deeply understand the meaning of sign language gestures and convert them into corresponding semantic information, that is, meaning that can be described in words or language.

[0123] Step S730: Generate audio information based on the semantic information.

[0124] In this embodiment of the disclosure, after the sign language information is successfully converted into semantic information, this information can be further converted into audio information. This means that the originally silent sign language actions are now given sound and become audible language.

[0125] In an exemplary embodiment, the audio information may be speech synthesized by a text-to-speech (TTS) system or a pre-recorded audio segment that matches the semantic information.

[0126] Step S740: Play the audio information.

[0127] In this embodiment of the disclosure, the generated audio information can be played through a speaker or other audio output device, allowing sign language users and / or other listeners to hear the content expressed by the sign language information. The speaker or other audio output device can be the same as or different from the device implementing the speech conversion function. When the speaker or other audio output device is different from the device implementing the speech conversion function, audio information and control signals can be sent to the speaker or other audio output device to cause it to play the audio information.

[0128] For example, in special education schools, teachers can use this function to better understand the sign language of deaf students, while students can confirm the teacher's instructions through auditory feedback, thereby improving teaching effectiveness. As another example, in public places such as hospitals and banks, deaf individuals can interact with the system through sign language and receive instant voice feedback, making services more convenient. Furthermore, deaf individuals can control smart home devices at home using sign language and hear confirmation messages from the system, enjoying a more intelligent living experience.

[0129] By incorporating a speech-to-audio conversion function to process sign language information and convert it into audio for playback, this disclosure provides a more convenient and efficient communication method for people with language impairments. This design is not only applicable to the field of accessible communication but can also be widely applied in education, training, entertainment, and other fields, providing more intelligent and user-friendly services for users with diverse needs.

[0130] Figure 8 This is a schematic diagram of a system for implementing a gesture recognition method, according to some embodiments of this disclosure. Figure 8 As shown, it includes a camera, an event camera, and processing circuitry. The camera may include an image processing unit and a low-power graphics processing unit. The processing circuitry may include multiple modules with different functions, such as a low-power image processing system, an image processing system, a video encoder, and a display module.

[0131] refer to Figure 8 The low-power graphics processing unit and event camera in the camera can be used to monitor and capture specified gestures in a low-power image acquisition mode. Users can customize different gestures to activate different gesture recognition functions. For example, gesture 1 (such as a clenched fist gesture) can be defined to activate the sign language translation function, and gesture 2 (such as a heart gesture) can be defined to activate the command conversion function.

[0132] Then, a low-power image processing system can be invoked to process the specified gesture and determine the gesture recognition function corresponding to the specified gesture. For example, when the comparison and detection is a fist gesture 1, it can be determined that the sign language translation function needs to be activated at this time.

[0133] Gesture recognition can be achieved using image processing systems and video encoders, with input from gesture information captured by cameras and event cameras. Furthermore, distributed computing power deployment for sign language recognition can be achieved by combining edge systems and wireless computing centers.

[0134] like Figure 8 The system shown can perform the following operations depending on the gesture recognition function:

[0135] (1) Sign language translation function: Receive sign language information, obtain translated text information through translation algorithm, and display it through display module; at the same time, generate multiple fuzzy recommendation text information based on the translated text information, and display it to the user for selection through display module.

[0136] (2) Command conversion function: Receive gesture information, call the command library for matching, convert the gesture information into gesture commands, control the corresponding device to execute the command, and display the execution result; for sign language information, it is also necessary to first parse the sign language information to determine the target and processing method, and then generate gesture commands.

[0137] (3) Voice conversion function: Receive sign language information, obtain the corresponding audio information through a speech synthesis algorithm, and play the audio information.

[0138] It should be noted that the above figures are merely illustrative representations of the processes included in methods according to some embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0139] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the embodiments of the method disclosed herein.

[0140] Figure 9 This is a block diagram illustrating a gesture recognition device 900 according to some embodiments of the present disclosure. (Refer to...) Figure 9 The device includes a wake-up unit 901 and a processing unit 902.

[0141] The wake-up unit 901 is used to wake up the gesture recognition function corresponding to the specified gesture in response to the capture of the specified gesture; the processing unit 902 is used to process the gesture information through the gesture recognition function in response to the receipt of gesture information and display the processing result.

[0142] In some embodiments of this disclosure, the wake-up unit 901 is further configured to: capture the specified gesture through a first image acquisition mode; the processing unit 902 is further configured to: receive the gesture information through a second image acquisition mode in response to the gesture recognition function being woken up; wherein the power consumption of the first image acquisition mode is lower than the power consumption of the second image acquisition mode.

[0143] In some embodiments of this disclosure, the gesture recognition function includes a sign language translation function; the gesture information is sign language information; wherein, in response to receiving the gesture information, the processing unit 902 processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving the sign language information, translating the sign language information through the sign language translation function to obtain translated text information; and displaying and / or broadcasting the translated text information.

[0144] In some embodiments of this disclosure, the processing unit 902 is further configured to: translate the sign language information using the sign language translation function to obtain initial text information; generate multiple fuzzy recommended text information based on the initial text information; display the multiple fuzzy recommended text information; receive a selection operation on the multiple fuzzy recommended text information; and determine the translated text information based on the selected fuzzy recommended text information.

[0145] In some embodiments of this disclosure, the gesture recognition function includes an instruction conversion function; wherein, in response to receiving gesture information, the processing unit 902 processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving gesture information, converting the gesture information into a gesture instruction through the instruction conversion function; controlling the object of the gesture instruction to execute the gesture instruction; and displaying the execution result of the gesture instruction.

[0146] In some embodiments of this disclosure, the gesture information is sign language information; wherein, in response to receiving the gesture information, the processing unit 902 converts the gesture information into a gesture command through the command conversion function, including: in response to receiving the sign language information, parsing the sign language information through the command conversion function to determine the target and processing method indicated by the sign language information; and generating the gesture command according to the target and processing method.

[0147] In some embodiments of this disclosure, the processing unit 902, in response to receiving gesture information, converts the gesture information into a gesture command through the command conversion function, including: in response to receiving gesture information, calling the command library associated with the command conversion function; and in response to the gesture information matching a target preset command in the command library, determining the preset command as the gesture command.

[0148] In some embodiments of this disclosure, the gesture recognition function includes a speech conversion function; the gesture information is sign language information; wherein, in response to receiving gesture information, the processing unit 902 processes the gesture information through the gesture recognition function and displays the processing result, including: in response to receiving sign language information, processing the sign language information through the speech conversion function to obtain semantic information corresponding to the sign language information; generating audio information based on the semantic information; and playing the audio information.

[0149] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0150] Figure 10 This is a block diagram illustrating a device 1000 for gesture recognition according to some embodiments of the present disclosure. For example, device 1000 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.

[0151] Reference Figure 10 The device 1000 may include one or more of the following components: a processing component 1002, a memory 1004, a power component 1006, a multimedia component 1008, an audio component 1010, an input / output (I / O) interface 1012, a sensor component 1014, and a communication component 1016.

[0152] Processing component 1002 typically controls the overall operation of device 1000, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 1002 may include one or more processors 1020 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 1002 may include one or more modules to facilitate interaction between processing component 1002 and other components. For example, processing component 1002 may include a multimedia module to facilitate interaction between multimedia component 1008 and processing component 1002.

[0153] Memory 1004 is configured to store various types of data to support the operation of device 1000. Examples of this data include instructions for any application or method operating on device 1000, contact data, phonebook data, messages, pictures, videos, etc. Memory 1004 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0154] The power supply component 1006 provides power to the various components of the device 1000. The power supply component 1006 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 1000.

[0155] Multimedia component 1008 includes a screen that provides an output interface between the device 1000 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 1008 includes a front-facing camera and / or a rear-facing camera. When the device 1000 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0156] Audio component 1010 is configured to output and / or input audio signals. For example, audio component 1010 includes a microphone (MIC) configured to receive external audio signals when device 1000 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 1004 or transmitted via communication component 1016. In some embodiments, audio component 1010 also includes a speaker for outputting audio signals.

[0157] I / O interface 1012 provides an interface between processing component 1002 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0158] Sensor assembly 1014 includes one or more sensors for providing state assessments of various aspects of device 1000. For example, sensor assembly 1014 may detect the on / off state of device 1000, the relative positioning of components such as the display and keypad of device 1000, changes in the position of device 1000 or a component of device 1000, the presence or absence of user contact with device 1000, the orientation or acceleration / deceleration of device 1000, and temperature changes of device 1000. Sensor assembly 1014 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1014 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1014 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0159] Communication component 1016 is configured to facilitate wired or wireless communication between device 1000 and other devices. Device 1000 can access wireless networks based on communication standards, such as WiFi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 1016 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 1016 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0160] In some embodiments of this disclosure, the apparatus 1000 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0161] In some embodiments of this disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1004 including instructions, which can be executed by a processor 1020 of the device 1000 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0162] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of a mobile terminal, enables the mobile terminal to execute a gesture recognition method, the method comprising: activating a gesture recognition function corresponding to the specified gesture in response to capturing a specified gesture; and processing the gesture information through the gesture recognition function in response to receiving gesture information, and displaying the processing result.

[0163] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0164] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A gesture recognition method, characterized in that, include: In response to the capture of a specified gesture, the gesture recognition function corresponding to the specified gesture is activated; In response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result.

2. The method according to claim 1, characterized in that, The method further includes: The specified gesture is captured using the first image acquisition mode; In response to the activation of the gesture recognition function, the gesture information is received through a second image acquisition mode; wherein the power consumption of the first image acquisition mode is lower than that of the second image acquisition mode.

3. The method according to claim 1 or 2, characterized in that, The gesture recognition function includes a sign language translation function; the gesture information is sign language information. Specifically, in response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result, including: In response to receiving sign language information, the sign language information is translated using the sign language translation function to obtain translated text information; Display and / or broadcast the translated text information.

4. The method according to claim 3, characterized in that, The sign language information is translated using the sign language translation function to obtain translated text information, including: The sign language information is translated using the sign language translation function to obtain initial text information; Multiple fuzzy recommendation text messages are generated based on the initial text information; Display the multiple fuzzy recommendation text information; The system receives a selection operation on the plurality of fuzzy recommended text information and determines the translated text information based on the selected fuzzy recommended text information.

5. The method according to claim 1 or 2, characterized in that, The gesture recognition function includes a command conversion function; Specifically, in response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result, including: In response to receiving gesture information, the gesture information is converted into gesture commands through the command conversion function; The object to which the gesture command is applied executes the gesture command; The execution result of the gesture command is displayed.

6. The method according to claim 5, characterized in that, The gesture information is sign language information; Specifically, in response to receiving gesture information, the gesture information is converted into gesture commands through the command conversion function, including: In response to receiving sign language information, the sign language information is parsed through the instruction conversion function to determine the target and processing method indicated by the sign language information; The gesture command is generated based on the target object and processing method.

7. The method according to claim 5, characterized in that, In response to receiving gesture information, the gesture information is converted into gesture commands through the command conversion function, including: In response to receiving gesture information, the instruction library associated with the instruction conversion function is invoked; In response to the gesture information matching a target preset instruction in the instruction library, the preset instruction is determined as the gesture instruction.

8. A gesture recognition device, characterized in that, include: A wake-up unit is used to wake up the gesture recognition function corresponding to the specified gesture in response to the capture of the specified gesture; The processing unit is configured to respond to received gesture information by processing the gesture recognition function and displaying the processing result.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the steps of the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of a mobile terminal, enable the mobile terminal to perform a gesture recognition method, the method comprising: In response to the capture of a specified gesture, the gesture recognition function corresponding to the specified gesture is activated; In response to receiving gesture information, the gesture recognition function processes the gesture information and displays the processing result.