A gesture control method, vehicle, device, storage medium, and program product

CN122569723APending Publication Date: 2026-08-14HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]目前,车机等电子设备在通过神经网络模型实现车辆的手势控制功能时,会占用电子设备较多的计算资源和存储资源

Benefits of technology

[0044] Fifthly, this application provides a computer program product, including: computer instructions, which, when executed on an electronic device, cause the electronic device to perform the gesture control method mentioned in this application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569723A_ABST
    Figure CN122569723A_ABST
Patent Text Reader

Abstract

This application relates to the field of device control technology, and discloses a gesture control method, vehicle, device, storage medium, and program product. The gesture control method includes: after acquiring a video of a user's dynamic gestures recorded by a camera, the electronic device can determine the gesture features to be processed corresponding to the user's gesture operation based on the positions of key points of the user's hand (e.g., finger joints, fingertips, etc.) in each frame of the video. Next, the electronic device can compare the gesture features to be processed with pre-stored template gesture features. When the electronic device determines that the similarity between the user's gesture features to be processed and a specific template gesture feature is high, it can execute the control operation corresponding to that specific template gesture feature. In this way, the electronic device can improve the accuracy of user gesture operation classification while reducing the consumption of computing and storage resources, thereby accurately realizing the gesture control function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of equipment control technology, and in particular to a gesture control method, vehicle, equipment, storage medium, and program product. Background Technology

[0002] With the development of intelligent technology, more and more vehicles are equipped with gesture control functions, allowing users to directly control in-vehicle devices remotely through gestures, thus maximizing user convenience in operating and using these devices. For example... Figure 1 As shown, when a user is sitting in the driver's seat of vehicle 101, if they need to open or close the passenger-side door, they only need to make a corresponding gesture (e.g., wave their hand) from the driver's seat. The vehicle's infotainment system will then respond to the gesture to control the opening or closing of the passenger-side door. Thus, through the gesture control function inside vehicle 101, users do not need to get out of the car or reach out to operate the door, providing a more comprehensive and personalized driving experience.

[0003] Currently, when in-vehicle infotainment systems and other electronic devices implement gesture control functions using neural network models, they consume significant computing and storage resources. Furthermore, in practical engineering applications, the limited computing power or memory of electronic devices may prevent the loading and execution of all neural network layers in the model, thus affecting the usability of gesture control functions. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a gesture control method, vehicle, device, storage medium, and program product.

[0005] In a first aspect, this application provides a gesture control method applied to an electronic device. The method includes: acquiring a first dynamic gesture video to be processed; determining a first gesture feature based on the position of each hand key point in each video frame of the first dynamic gesture video; determining the similarity between the first gesture feature and each preset template gesture feature; and executing a first control operation corresponding to the first template gesture feature when the similarity between the first gesture feature and the preset first template gesture feature meets a similarity condition.

[0006] In this application, the first dynamic gesture video can be a gesture operation video recorded when a user performs a gesture operation as mentioned in this application; the first gesture feature can be a user gesture feature corresponding to the user gesture operation mentioned in this application; the first template gesture feature can be a specific template gesture feature mentioned in this application that satisfies the similarity condition with the first gesture feature; and the first control operation can be the control operation corresponding to the specific template gesture feature mentioned in this application.

[0007] In some implementations, recording the first dynamic gesture video of a user performing a gesture operation via a camera can capture any type of complete user gesture operation and realize the corresponding gesture control function.

[0008] In some implementations, template gesture features corresponding to each control operation can be pre-stored in the electronic device. When the electronic device determines the first gesture feature corresponding to the user's gesture operation based on the positions of key points of the user's hand (e.g., finger joints, fingertips, etc.) in each frame of video, the electronic device can compare the first gesture feature with each pre-stored template gesture feature. Specifically, when the electronic device determines that the first gesture feature has a high similarity to a specific first template gesture feature, it can execute the first control operation corresponding to that first template gesture feature. For example, when the electronic device determines that the user's first gesture feature A has a high similarity to a pre-stored first template gesture feature A′ for closing the passenger side door, the electronic device can execute the control operation corresponding to that first template gesture feature A′ to close the passenger side door, thereby closing the passenger side door.

[0009] Thus, by using the above method, when classifying user gesture operations, it is only necessary to calculate the similarity between the gesture feature to be processed corresponding to the user gesture operation and the gesture features of each pre-stored template. That is, the method provided in this application embodiment does not require the use of a classification neural network model that consumes a large amount of memory and computing resources to implement gesture control functions, effectively reducing the consumption of computing and storage resources of electronic devices.

[0010] In one possible implementation of the first aspect above, the type of the first gesture feature includes at least one of feature vector, matrix, and feature map; and the type of the first gesture feature is the same as the type of each template gesture feature.

[0011] In some implementations, the first gesture feature corresponding to the user's gesture operation is of the same type as the pre-stored template gesture features, and the type of the first gesture feature can be arbitrary; this application does not limit this. For example, when the type of the first gesture feature and each template gesture feature is a feature vector, the electronic device can calculate the similarity between the first gesture feature vector and each template feature vector in the preset template feature vector library, and execute the control operation corresponding to that first template feature vector when the similarity between the first gesture feature vector and a certain first template feature vector meets the similarity condition. As another example, when the type of the first gesture feature and each template gesture feature is a matrix, the electronic device can calculate the similarity between the user gesture feature matrix and each template feature matrix, and execute the control operation corresponding to that specific template feature matrix when the similarity between the user gesture feature matrix and a certain specific template feature matrix meets the similarity condition.

[0012] Thus, when the first gesture feature corresponding to the user's gesture operation is of the same type as the pre-stored template gesture features, it is beneficial for the electronic device to calculate the similarity between the first gesture feature and each template gesture feature.

[0013] In one possible implementation of the first aspect above, the type corresponding to the first gesture feature and each template gesture feature is a feature vector. Determining the first gesture feature based on the position of each hand key point in each video frame of the first dynamic gesture video includes: determining the first gesture feature vector based on the position of each hand key point in each video frame of the first dynamic gesture video; determining the similarity between the first gesture feature and each preset template gesture feature includes: determining the similarity between the first gesture feature vector and each template feature vector in the preset template feature vector library; and, when the similarity between the first gesture feature and the preset first template gesture feature satisfies a similarity condition, executing the first control operation corresponding to the first template gesture feature includes: executing the first control operation corresponding to the first template feature vector when the similarity between the first gesture feature vector and the first template feature vector satisfies a similarity condition.

[0014] In this application, the first gesture feature vector can be the first gesture feature mentioned in this application, represented in the form of a feature vector; the first template feature vector can be the template feature vector mentioned in this application, whose similarity to the first gesture feature vector satisfies the similarity condition.

[0015] In some implementations, when both the first gesture feature and each template gesture feature are feature vectors, a template feature vector library can be pre-set in the electronic device, storing multiple template feature vectors and the corresponding control operations. The electronic device can determine the first gesture feature vector based on the position of each hand key point in multiple video frames of the first dynamic gesture video. Next, the electronic device can determine the similarity between the first gesture feature vector and each template feature vector in the pre-set template feature vector library to determine the first template feature vector that meets the similarity condition, thereby executing the first control operation corresponding to the first template feature vector.

[0016] Thus, the above gesture control method can improve the accuracy of user gesture operation classification, thereby accurately realizing gesture control function. At the same time, the above method can also effectively reduce the consumption of computing and storage resources of electronic devices.

[0017] In one possible implementation of the first aspect described above, the dimensions of each template feature vector in the preset template feature vector library are the same; and the dimension of the first gesture feature vector is the same as the dimension of each template feature vector.

[0018] In some implementations, if the dimension of each template feature vector is m, then the dimension of the user gesture feature vector is also m. This facilitates the electronic device in calculating the similarity between the user gesture feature vector and each template feature vector.

[0019] In one possible implementation of the first aspect above, the method of obtaining each feature data in the first gesture feature vector includes: determining the gesture feature data group corresponding to each video frame based on the position of each hand key point in each video frame, and obtaining multiple gesture feature data groups; obtaining the average value between the corresponding feature data in the multiple gesture feature data groups, and obtaining each feature data in the first gesture feature vector.

[0020] For example, if there is a one-to-one correspondence between A1 in the m feature data corresponding to the first video frame, A2 in the m feature data corresponding to the second video frame, A3 in the m feature data corresponding to the third video frame, and An in the m feature data corresponding to the nth video frame, then the electronic device can record the average value (A1+A2+A3...+An) / n among the one-to-one correspondence data in the first gesture feature vector. Similarly, by calculating the average value of the corresponding feature data in the feature data group of each video frame (each feature data group may include m feature data), the m feature data in the first gesture feature vector can be obtained. In this way, the dimension of the user gesture feature vector corresponding to each dynamic gesture video can be kept consistent.

[0021] In one possible implementation of the first aspect described above, the feature components in the first gesture feature vector may further include at least one of a static posture feature component, a displacement feature component, and a hand orientation feature component; wherein, the static posture feature component is used to indicate the hand posture, which includes at least one of the number of extended fingers and the shape fitted by the hand; the displacement feature component is used to indicate the displacement of each hand key point in the first dynamic gesture video; and the hand orientation feature component is used to indicate the orientation of the palm.

[0022] Thus, through the aforementioned multiple feature components, the user's gesture operation in the first dynamic gesture video can be described in more detail.

[0023] The acquisition method for each feature data in the static pose feature portion or hand orientation feature portion can include: determining the gesture feature data group (e.g., static pose feature data group or hand orientation feature data group) corresponding to each video frame based on the position of each hand key point in multiple consecutive frames over a period of time, thereby obtaining multiple gesture feature data groups. Next, the electronic device can obtain the average value between the corresponding feature data in the multiple gesture feature data groups to obtain each feature data in the static pose feature portion or hand orientation feature portion. For example, over a period of time spanning multiple consecutive frames (e.g., 3 frames), the electronic device determines the static pose feature data in the first video frame as (e.g., L1', L2'); the static pose feature data in the second video frame as (L1'', L2''); and the static pose feature data in the third video frame as (L1''', L2'''). The electronic device can then calculate the average value of the corresponding static pose feature data in the multiple static pose feature data groups to obtain the static pose feature components (L1, L2).

[0024] Furthermore, the acquisition methods for the feature data in the displacement feature section can include: determining the displacement of each hand keypoint over a period of time based on its position across multiple consecutive frames, thereby obtaining the feature data in the displacement feature section. For example, taking two video frames as an example, if the coordinates of hand keypoint A in the first video frame are p1 and its coordinates in the second video frame are p2, then the electronic device can determine the displacement information of hand keypoint A in the two video frames based on p1 and p2. Similarly, the electronic device can also determine the displacement information of other hand keypoints, thus obtaining the feature data in the displacement feature section.

[0025] In one possible implementation of the first aspect described above, the first gesture feature vector may further include n sub-sequence feature vectors; and the determination of the first gesture feature vector based on the position of each hand key point in each video frame of the first dynamic gesture video further includes: obtaining n video segments based on a preset sliding window, wherein the length of the sliding window is x1 video frames, the sliding step size of the sliding window is x2 video frames, and x1 and x2 are both positive integers; and determining the sub-sequence feature vector corresponding to each video segment in the n video segments based on the position of each hand key point in each video segment in each frequency frame.

[0026] In a possible implementation of the above first aspect, obtaining n video segments based on a preset sliding window includes: taking the video segment between the i-th video frame and the (i + x1 - 1)-th video frame corresponding to the sliding window as the h-th video segment; moving the sliding window by x2 video frames, and taking the video segment between the (i + x2)-th video frame and the (i + x2 + x1 - 1)-th video frame corresponding to the sliding window as the (h + 1)-th video segment; where i ≥ 1 and i is a positive integer, 1 ≤ h < n and h is a positive integer.

[0027] For example, in some embodiments, if the length x1 is 10 and the sliding step x2 is 2, the electronic device may take the video segment between the 1st frame and the 10th frame (i.e., i to i + x1 - 1) in the first dynamic gesture video as the first video segment. Then, the electronic device can control the sliding window to move a distance of 2 video frames, and take the video segment between the 3rd frame and the 12th frame as the second video segment; next, the electronic device can control the sliding window to move a distance of 2 video frames again, and take the video segment between the 5th frame and the 14th frame as the third video segment. Similarly, the electronic device can control the sliding window to move a distance of 2 video frames in sequence, so as to obtain n video segments.

[0028] In a possible implementation of the above first aspect, determining the similarity between the first gesture feature vector and each template feature vector in the preset template feature vector library includes: calculating the subsequence similarity between each subsequence feature vector in the n subsequence feature vectors and each template feature vector respectively, to obtain n subsequence similarities between the first gesture feature vector and each template feature vector.

[0029] In a possible implementation of the above first aspect, corresponding to the similarity between the first gesture feature vector and the first template feature vector satisfying the similarity condition, performing the first control operation corresponding to the first template feature vector includes: corresponding to k subsequence similarities among the n subsequence similarities between the first gesture feature vector and the first template feature vector satisfying the similarity condition, determining that the similarity between the first gesture feature vector and the first template feature vector satisfies the similarity condition, and performing the first control operation corresponding to the first template feature vector; where k ≤ n, and k is a positive integer greater than the preset quantity threshold.

[0030] In this way, electronic devices can classify gesture operations in each video segment. For example, if multiple video segments all contain gesture operation A, it can be determined that the first gesture feature vector corresponding to that user gesture operation satisfies the similarity condition with the first template feature vector corresponding to gesture operation A. Similarly, if only one or a very small number of video segments contain gesture operation A, while multiple consecutive video segments contain gesture operation B, it can be determined that the first gesture feature vector corresponding to that user gesture operation satisfies the similarity condition with the first template feature vector corresponding to gesture operation B. Therefore, classifying user gesture operations based on the sub-sequence feature vectors corresponding to multiple video segments can make the classification results more accurate.

[0031] In one possible implementation of the first aspect above, determining the similarity between the first gesture feature vector and each template feature vector in the preset template feature vector library includes: calculating the Euclidean distance between the first gesture feature vector and each template feature vector; using each Euclidean distance as the similarity between the first gesture feature vector and the corresponding template feature vector; and the similarity between the first gesture feature vector and the first template feature vector satisfies the similarity condition, including: the Euclidean distance between the first gesture feature vector and the first template feature vector is less than a preset distance threshold.

[0032] In some implementations, the electronic device can calculate the similarity between the first gesture feature vector and each template feature vector in a variety of ways, such as Euclidean distance, cosine similarity, Manhattan distance, etc., which are not limited in this application.

[0033] In one possible implementation of the first aspect above, the electronic device may further include a custom mode, and the method further includes: when the electronic device is in custom mode, acquiring a second dynamic gesture video; determining a second gesture feature vector corresponding to the second dynamic gesture video based on the position of each hand key point in each video frame of the second dynamic gesture video; adding the second gesture feature vector to a preset template feature vector library; acquiring the user's operation of associating the second gesture feature vector with a second control operation, and storing the association relationship between the second gesture feature vector and the second control operation.

[0034] In some implementations, when the electronic device is in custom mode, the user can choose to bind a new second control operation to the second gesture feature vector, or replace other template feature vectors corresponding to the second control operation with the second gesture feature vector. Thus, the above method also allows the user to maintain a preset template feature vector library to achieve scalability of the template feature vector library.

[0035] In one possible implementation of the first aspect above, the method further includes: when the electronic device is in a custom mode, acquiring multiple third dynamic gesture videos; determining each third gesture feature vector corresponding to each third dynamic gesture video based on the position of each hand key point in each video frame of each third dynamic gesture video; adding the average feature vector between each third gesture feature vector corresponding to each third dynamic gesture video to a preset template feature vector library; acquiring the user's operation of associating the average feature vector with a third control operation, and storing the association relationship between the average feature vector and the third control operation.

[0036] This allows electronic devices to implement gesture control functions while also enabling users to maintain their own preset template feature vector library to achieve scalability of the template feature vector library.

[0037] In one possible implementation of the first aspect above, the execution of the first control operation corresponding to the first gesture feature and the preset first template gesture feature satisfying the similarity condition further includes: the electronic device sending an instruction corresponding to the first control operation to the execution device corresponding to the first control operation, corresponding to the first gesture feature and the first template gesture feature satisfying the similarity condition; wherein the instruction corresponding to the first control operation is used to cause the execution device to perform the first control operation.

[0038] In some implementations, if the executing device is an in-vehicle device, the electronic device (e.g., a vehicle infotainment system) can send the instruction corresponding to the first control operation to the in-vehicle device, causing the in-vehicle device to perform the first control operation; as another example, if the executing device is a smart home device, the electronic device (e.g., a smart terminal) can send the instruction corresponding to the first control operation to the smart home device, causing the smart home device to perform the first control operation.

[0039] Thus, the above gesture control method can improve the accuracy of user gesture operation classification, thereby accurately realizing gesture control function. At the same time, the above method can also effectively reduce the consumption of computing and storage resources of electronic devices.

[0040] Secondly, this application provides a vehicle, including: a camera for capturing dynamic gesture video; a vehicle-mounted system for determining a first gesture feature based on the position of each hand key point in each video frame of the dynamic gesture video; and the vehicle-mounted system is further configured to calculate the similarity between the first gesture feature and each preset template gesture feature, and when it is determined that the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition, the vehicle-mounted system is further configured to control an execution device in the vehicle corresponding to the first template gesture feature to perform a corresponding first control operation; the execution device is configured to perform the first control operation in response to the control command of the vehicle-mounted system.

[0041] In some embodiments, the gesture control method provided in this application can be applied to any scenario where gesture-based air control can be achieved. Specifically, when the gesture control method provided in this application is used to control in-vehicle devices, the vehicle may include a camera, a vehicle infotainment system, and execution devices. Thus, through the interaction between various modules within the vehicle, the vehicle can implement the gesture control method mentioned in this application.

[0042] Thirdly, this application provides an electronic device, including: a memory and a processor, wherein the memory is coupled to the processor; the memory is used to store computer program code / instructions; when the computer program code / instructions are executed by the processor, the electronic device performs the gesture control method mentioned in this application.

[0043] Fourthly, this application provides a readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the gesture control method mentioned in this application.

[0044] Fifthly, this application provides a computer program product, including: computer instructions, which, when executed on an electronic device, cause the electronic device to perform the gesture control method mentioned in this application.

[0045] The beneficial effects of the second to fifth aspects mentioned above can be referred to the relevant descriptions in the first aspect and its various possible implementations, which will not be repeated here. Attached Figure Description

[0046] Figure 1 According to some embodiments of this application, a schematic diagram of an application scenario for remotely controlling in-vehicle devices through gestures is shown;

[0047] Figure 2 According to some embodiments of this application, a schematic diagram of an application scenario for determining user gesture feature vectors based on key points is shown;

[0048] Figure 3A According to some embodiments of this application, a schematic diagram of a first application scenario of gesture control operation is shown;

[0049] Figure 3B According to some embodiments of this application, a schematic diagram of a scenario in which gesture control function is enabled in a mobile phone is shown;

[0050] Figure 4 According to some embodiments of this application, a schematic diagram of a second application scenario for gesture control operation is shown;

[0051] Figure 5AAccording to some embodiments of this application, a schematic diagram of a third application scenario for gesture control operation is shown;

[0052] Figure 5B According to some embodiments of this application, a schematic diagram of a scenario in which gesture control function is activated in a vehicle is shown;

[0053] Figure 6 According to some embodiments of this application, a flowchart of a gesture control method is shown;

[0054] Figure 7 According to some embodiments of this application, a detailed flowchart of a gesture control method is shown;

[0055] Figure 8 According to some embodiments of this application, a schematic diagram of a scenario for controlling in-vehicle equipment based on a gesture control method is shown;

[0056] Figure 9 According to some embodiments of this application, a flowchart of a custom mode is shown;

[0057] Figure 10 According to some embodiments of this application, a schematic diagram of a scenario in which a custom mode is executed on a vehicle is shown;

[0058] Figure 11 According to some embodiments of this application, a schematic diagram of the processing flow of various modules of a vehicle for a gesture control method is shown;

[0059] Figure 12 According to some embodiments of this application, a schematic diagram of the hardware structure of an electronic device is shown. Detailed Implementation

[0060] The illustrative embodiments of this application include, but are not limited to, a gesture control method, a vehicle, a device, a storage medium, and a program product.

[0061] As mentioned earlier, the trained deep learning-based neural network model can be directly deployed in the vehicle's infotainment system, enabling the system to implement gesture control functions. The infotainment system is an embedded device running an operating system, through which the vehicle can perform navigation, entertainment, and communication functions.

[0062] Furthermore, neural network models typically include multiple neural network layers, such as pooling layers to reduce feature dimensionality and improve model generalization ability (i.e., the adaptability of deep learning models to new input data), and fully connected layers to integrate features and perform feature recognition. Each neural network layer contains multiple neurons, and multiple parameters or functions are required between connected neurons. For example, parameters between two neurons may include weights to determine the influence of the previous neuron's output on the current neuron; biases to determine under what conditions the current neuron will be activated; and activation functions to avoid directly using the previous neuron's output as the current neuron's input (i.e., avoiding linear processing). It is evident that neural network models contain a large number of parameters and functions, causing electronic devices such as in-vehicle systems to consume significant computational and storage resources when implementing gesture control functions through neural network models.

[0063] It is understood that user gesture operations used to implement gesture control functions can include, but are not limited to, any of the following gesture operations: making gestures such as "OK" or "V"; waving to the left / right, bending the hand downwards, raising the hand upwards, etc.; and gesture operations that continuously switch between gestures such as "OK" and / or waving actions. Therefore, this application can record user gesture operations via a camera, thereby capturing any type of complete user gesture operation in the recorded user gesture operation video (which can be described as "first dynamic gesture video") and implementing the corresponding gesture control function.

[0064] In the gesture control method provided in this application, after acquiring the first dynamic gesture video recorded by the camera, the electronic device can determine the gesture feature to be processed corresponding to the user's gesture operation (as an example of the first gesture feature) based on the position of each key point of the user's hand (e.g., finger joints, fingertips, etc.) in each frame of the video. Next, the electronic device can compare the gesture feature to be processed with each pre-stored template gesture feature. When the electronic device determines that the user's gesture feature to be processed has a high similarity to a specific template gesture feature (as an example of the first template gesture feature), it can execute the control operation corresponding to that specific template gesture feature (as an example of the first control operation). For example, when the electronic device determines that the user's gesture feature A to be processed has a high similarity to a pre-stored specific template gesture feature A′ for closing the passenger side door, the electronic device can execute the control operation corresponding to that specific template gesture feature A′ to close the passenger side door, thereby closing the passenger side door.

[0065] Thus, by using the above method, when classifying user gesture operations, it is only necessary to calculate the similarity between the gesture feature to be processed corresponding to the user gesture operation and the gesture features of each pre-stored template. That is, the method provided in this application embodiment does not require the use of a classification neural network model that consumes a large amount of memory and computing resources to implement gesture control functions, effectively reducing the consumption of computing and storage resources of electronic devices.

[0066] It is understood that the type of the gesture feature to be processed corresponding to the user's gesture operation is the same as that of the pre-stored template gesture features, and the type of the gesture feature to be processed can be arbitrary; this application does not limit this. For example, when the type of the gesture feature to be processed and the template gesture features is a feature vector, such as the user gesture feature vector corresponding to the user's gesture operation being (gesture feature data 1, gesture feature data 2, gesture feature data 3, gesture feature data 4...), the electronic device can calculate the similarity between the user gesture feature vector and each template feature vector in the preset template feature vector library, and execute the control operation corresponding to that specific template feature vector when the similarity between the user gesture feature vector and a specific template feature vector meets the similarity condition. As another example, when the type of the gesture feature to be processed and the template gesture features is a matrix, such as the user gesture feature matrix corresponding to the user's gesture operation being... The electronic device can calculate the similarity between the user's gesture feature matrix and each template feature matrix, and execute the control operation corresponding to that specific template feature matrix when the similarity between the user's gesture feature matrix and a specific template feature matrix meets the similarity condition. Furthermore, the type of gesture feature to be processed can also be a feature map, where each gesture feature data is represented by pixel data in the feature map. Therefore, this application does not limit the type of gesture feature to be processed.

[0067] For ease of description, in the following embodiments, the gesture control method provided in this application is described using the example that the gesture features to be processed corresponding to the user's gesture operation and the gesture features of each pre-stored template are both feature vectors.

[0068] In this embodiment of the application, the electronic device can determine the user's gesture feature vector (as an example of a first gesture feature vector) based on the position of the user's hand key points in each video frame. For example, Figure 2 Figure (a) illustrates a gesture operation where five fingers are extended and the palm faces inward towards the screen. Figure 2 Figure (b) illustrates a gesture involving extending one finger with the palm facing inwards towards the screen. The black dots represent key hand points detected by the electronic device. (Comparison) Figure 2As shown in Figures (a) and (b), the position of each hand key point, or the distance between specific hand key points, varies depending on the gesture operation. Therefore, this application can determine the user's gesture feature vector based on the position of the user's hand key points in each video frame.

[0069] In some embodiments, for ease of management and storage, the dimensions (which can be the number of feature data points) of each template feature vector in the preset template feature vector library can be the same. For example, if the dimension of each template feature vector is m (i.e., including m feature data points), then each template feature vector can be (B1, B2, B3...Bm). It should be understood that when comparing the user gesture feature vector (as an example of a first gesture feature vector) with each template feature vector, the dimension of the user gesture feature vector should also be the same as the dimension of each template feature vector. For example, if the dimension of each template feature vector is m, then the dimension of the user gesture feature vector is also m.

[0070] For example, in some embodiments, the electronic device can determine the gesture feature data group corresponding to each video frame based on the position of each hand key point in each video frame. For example, it can determine m feature data corresponding to the first video frame, m feature data corresponding to the second video frame, m feature data corresponding to the third video frame, and so on. Then, the electronic device can perform statistics or calculate the average value of the corresponding feature data in multiple feature data groups to obtain each feature data in the first gesture feature vector. For example, if A1 in the m feature data corresponding to the first video frame, A2 in the m feature data corresponding to the second video frame, A3 in the m feature data corresponding to the third video frame, and An in the m feature data corresponding to the nth video frame are one-to-one, the electronic device can record the average value (A1+A2+A3...+An) / n between the one-to-one corresponding feature data in the first gesture feature vector. Similarly, by calculating the average value of the corresponding feature data in the feature data group (each feature data group includes m feature data) corresponding to each video frame, the m feature data in the first gesture feature vector can be obtained. In this way, the dimension of the user gesture feature vector corresponding to each dynamic gesture video can be kept consistent.

[0071] Furthermore, in other embodiments, the user gesture feature vector can also be composed of multiple feature components. For example, when the user gesture feature vector is composed of multiple feature components such as static posture (e.g., the number of extended fingers), hand orientation (e.g., palm facing inwards from the screen), and hand displacement (e.g., the displacement of key points in the first dynamic gesture video), the pre-stored template feature vectors should also be composed of the aforementioned multiple feature components, thereby facilitating the electronic device to calculate the similarity between the user gesture feature vector and each template feature vector. For example, if the representation of each template feature vector is ([static posture 1′, static posture 2′…static posture n′], [hand orientation 1′, hand orientation 2′…hand orientation k′], [hand displacement 1′, hand displacement 2′…hand displacement q′]), then when the electronic device determines the user gesture feature vector represented by ([static posture 1, static posture 2…static posture n], [hand orientation 1, hand orientation 2…hand orientation k], [hand displacement 1, hand displacement 2…hand displacement q]), it can calculate the similarity between the user gesture feature vector and each template feature vector.

[0072] In summary, in this application, even though different users may perform the same gesture operation at different speeds, resulting in different durations of the dynamic gesture videos recorded by the camera, the dimension of the user gesture feature vector determined in this application is fixed. That is, the dimension of the user gesture feature vector is independent of the duration of the dynamic gesture video. Furthermore, in the embodiments of this application, the dimension, representation, and feature components of the first gesture feature vector are the same as those of the feature vectors of each template.

[0073] Furthermore, in some embodiments, the electronic device can calculate the similarity between the user gesture feature vector and each template feature vector in various ways, and this application does not limit this. For example, the electronic device can use the Euclidean distance between the user gesture feature vector and each template feature vector as the similarity between the user gesture feature vector and the corresponding template feature vector. For example, if the m-dimensional user gesture feature vector is (A1, A2, A3...Am), and there exists an m-dimensional template feature vector (B1, B2, B3...Bm), then the Euclidean distance between the user gesture feature vector and the template feature vector can be... In addition, electronic devices can also determine the similarity between user gesture feature vectors and corresponding template feature vectors by calculating cosine similarity, Manhattan distance, etc., but this application does not limit this.

[0074] Specifically, when an electronic device determines that the similarity between a user's dynamic gesture feature vector and a specific template feature vector (as an example of a first template feature vector) meets a similarity condition, it can consider the user's dynamic gesture feature vector to be the same as the specific template feature vector, and thus execute the control operation corresponding to the specific template feature vector. The similarity condition can be that the similarity is greater than a similarity threshold; alternatively, when the electronic device represents similarity using Euclidean distance, the similarity condition can also be that the Euclidean distance is less than a preset distance threshold; this application does not limit this.

[0075] Thus, by using the gesture control method provided in this application, when classifying the feature vectors corresponding to user gesture operations, it is only necessary to calculate the similarity between the gesture feature vectors corresponding to user gesture operations and the pre-stored template feature vectors. That is, the method provided in this application embodiment does not require the use of a classification neural network model that consumes a lot of memory and computing resources to implement the gesture control function. In this way, while realizing the gesture control function, it can also effectively reduce the consumption of computing and storage resources of electronic devices.

[0076] It is understood that the gesture control method described in this application can be applied to any electronic device. These electronic devices include, but are not limited to, mobile stations (MS) and mobile terminals (MT). For example, electronic devices can be in-vehicle systems, mobile phones, smart TVs, wearable devices, tablets, desktop computers, laptops, virtual reality (VR) devices, augmented reality (AR) devices, terminals in industrial control, terminals in self-driving, terminals in remote medical surgery, terminals in smart grids, terminals in transportation safety, terminals in smart cities, and smart terminals in smart homes. This application does not limit the specific form of the electronic device.

[0077] Furthermore, it can be understood that the gesture control method provided in this application can be applied to any scenario where gesture-based air control can be achieved. For example, this includes, but is not limited to: gesture-based air control of terminal devices such as mobile phones and computers, gesture-based air control of smart home devices, and gesture-based air control of in-vehicle devices, etc. The following is a combination of... Figures 3A to 5B The schematic diagram shown illustrates the application scenarios of the gesture control method provided in this application.

[0078] The following is a combination of... Figure 3A and Figure 3B This describes scenarios where gesture-based air control is implemented on terminal devices such as mobile phones and computers.

[0079] like Figure 3A As shown in Figure (a), during video playback, the camera of mobile phone 301 can capture dynamic gesture videos of the user. Next, mobile phone 301 can detect the positions of key hand points in each video frame, thereby determining the user gesture feature vector. Then, mobile phone 301 can compare this user gesture feature vector with each template feature vector in a preset template feature vector library. For example, it can calculate the Euclidean distance between the user gesture feature vector and each template feature vector, and classify the user gesture feature vector into a specific template feature vector that satisfies the similarity condition (e.g., the Euclidean distance is less than a preset distance threshold and the Euclidean distance is minimal) (as an example of a first template feature vector). Finally, mobile phone 301 can execute the control operation corresponding to this specific template feature vector. For example, if... Figure 3A If the specific template feature vector corresponding to the gesture operation shown in Figure (a) (e.g., waving to the left) is used to instruct the phone 301 to lower the volume, then the phone 301 can display the interface 302 shown in Figure (b). In interface 302, a volume adjustment control 303 is displayed, and the phone 302 can directly manipulate the volume adjustment control 303 to lower the volume. Alternatively, if Figure 3A The feature vector of the feature template corresponding to the gesture operation shown in Figure (a) is used to instruct the mobile phone 301 to lower the volume. The mobile phone 301 may also directly control the volume to lower the volume to meet the user's needs without displaying the interface 302 and the volume adjustment control 303. This application does not limit this.

[0080] It is understood that this application does not limit the runtime of the provided gesture control method. For example, the function corresponding to the gesture control method provided in this application can be the default function of mobile phone 301. That is, after mobile phone 301 enters the working state, mobile phone 301 begins to execute the gesture control method provided in this application. As another example, the function corresponding to the gesture control method provided in this application can be an optional function of mobile phone 301, such as, see reference... Figure 3B As shown in Figure (a), when the mobile phone 301 detects that the user has enabled the gesture control function 3041 in the settings interface 304, the mobile phone 301 can begin to execute the gesture control method provided in this application. For example, the function corresponding to the gesture control method provided in this application can also be a function for a specific application, for example, see [reference]. Figure 3BAs shown in Figure (b), when the mobile phone 301 detects that the user has enabled the gesture control function 3051 permission in the permission management interface 305 of the video application, the mobile phone 301 can only start executing the gesture control method provided in this application after entering the corresponding video application. Therefore, this application does not limit the execution time of the gesture control method.

[0081] Furthermore, it can be understood that the gesture control method provided in this application can be used not only to control mobile phones, but also to control any terminal device such as computers, tablets, desktop computers, laptops, VR devices, and AR devices. This application does not limit this.

[0082] Thus, the gesture control method provided in this application enables the terminal device to effectively reduce the consumption of computing and storage resources while implementing gesture control functions.

[0083] The following is combined Figure 4 The diagram illustrates a scenario where smart home devices can be controlled remotely via gestures.

[0084] It is understood that when installing a smart home system, a smart terminal for implementing the gesture control method of this application can also be installed simultaneously. For example, the smart terminal may include a camera for capturing video of the user's dynamic gestures, a processing device for processing the user's dynamic gestures, and a driver for driving the smart home devices.

[0085] For example, refer to Figure 4 As shown in Figure (a), when the camera in the smart terminal captures a video of a user's dynamic gestures (as an example of a first dynamic gesture video), the smart terminal can detect the positions of key hand points in each video frame, thereby determining the user's gesture feature vector. Then, the smart terminal can compare this user gesture feature vector with each template feature vector in a preset template feature vector library. For example, it can calculate the Euclidean distance between the user gesture feature vector and each template feature vector, and classify the user gesture feature vector into specific template feature vectors that satisfy similarity conditions (e.g., the Euclidean distance is less than a preset distance threshold and the Euclidean distance is minimum). Finally, the smart terminal can execute the control operation corresponding to this specific template feature vector. For example, if... Figure 4 The specific template feature vector corresponding to the gesture operation (e.g., two-finger pinch operation) shown in Figure (a) is used to indicate closing the curtain 401, then as shown in Figure (a), the specific template feature vector is used to indicate closing the curtain 401. Figure 4 As shown in Figure (b), the driving device of the smart terminal can drive the curtain 401 to close, thereby realizing the user's gesture control needs.

[0086] In other embodiments, when installing smart home devices, in addition to installing a smart terminal capable of implementing the gesture control method of this application, only a camera for capturing user dynamic gesture videos can be installed. The camera can then transmit the captured user dynamic gesture videos to a mobile phone or other terminal device in real time. After determining a specific template feature vector with a high similarity to the user's gesture feature vector, the mobile phone or other terminal device can control the smart home devices to perform corresponding control operations.

[0087] Furthermore, it can be understood that the gesture control method provided in this application can be used not only to control curtains, but also to control any smart home devices such as televisions, speakers, refrigerators, and projectors, and this application does not limit it in this regard.

[0088] Thus, the gesture control method provided in this application enables terminal devices to achieve gesture control functions while effectively reducing the consumption of computing and storage resources of electronic devices such as smart terminals.

[0089] The following is combined Figure 5A and Figure 5B The diagram illustrates a scenario where in-vehicle devices can be controlled remotely via gestures.

[0090] It is understood that vehicles are typically equipped with an in-vehicle infotainment system, which is an embedded device running an operating system. Vehicles can use this system to perform functions such as navigation, entertainment, and communication. Therefore, in this embodiment, the gesture control method provided in this application can be executed through the in-vehicle infotainment system to achieve gesture-based remote control of in-vehicle devices.

[0091] For example, refer to Figure 5A As shown in Figure (a), after the camera in vehicle 501 captures a video of the user's dynamic gestures (as an example of a first dynamic gesture video), the vehicle's infotainment system in vehicle 501 can detect the positions of key hand points in each video frame, thereby determining the user's gesture feature vector in the dynamic gesture video. Then, the infotainment system can compare this user gesture feature vector with each template feature vector in a preset template feature vector library. For example, it can calculate the Euclidean distance between the user gesture feature vector and each template feature vector, and classify the user gesture feature vector into specific template feature vectors that satisfy similarity conditions (e.g., the Euclidean distance is less than a preset distance threshold and the Euclidean distance is minimum). Finally, the infotainment system can control the in-vehicle equipment corresponding to this specific template feature vector to perform corresponding control operations. For example, if... Figure 5A The specific template feature vector corresponding to the gesture operation shown in Figure (a) (e.g., waving to the left) is used to indicate opening the passenger side door of vehicle 501. Then, as shown in Figure (a),... Figure 5AAs shown in Figure (b), the vehicle's infotainment system can control the opening of the passenger-side door, thereby fulfilling the user's gesture control needs.

[0092] It is understood that the execution time of the gesture control method provided in this application is not limited in vehicle 501. For example, the function corresponding to the gesture control method provided in this application can be the default function of vehicle 501. That is, after vehicle 501 is started, vehicle 501 begins to execute the gesture control method provided in this application. As another example, the function corresponding to the gesture control method provided in this application can be an optional function of vehicle 501. For example, see reference... Figure 5B As shown, when the vehicle system detects that the user has enabled the gesture control function 5031 in the settings interface 503 of the vehicle screen 502, the vehicle 501 can begin to execute the gesture control method provided in this application. Therefore, this application does not limit the execution time of the gesture control method.

[0093] Furthermore, it can be understood that, in a vehicle, the gesture control method provided in this application can be used not only to control the doors, but also to control any in-vehicle equipment such as air conditioning, windows, and music playback devices; this application does not limit this.

[0094] Thus, the gesture control method provided in this application enables in-vehicle devices to achieve gesture control functions while effectively reducing the computational and storage resource consumption of electronic devices such as vehicle infotainment systems.

[0095] In summary, the gesture control method provided in this application can be applied to any scenario where gesture-based air control can be achieved. For example, gesture-based air control of terminal devices, gesture-based air control of smart home devices, and gesture-based air control of in-vehicle devices, etc., are not limited in this application.

[0096] The following is combined with Figure 6 The flowchart shown describes the gesture control method provided in this application. It can be understood that in the embodiments of this application, Figure 6 The execution subject of each process in the method shown can be an electronic device, such as the aforementioned mobile phone 301, smart terminal, or in-vehicle system, etc. The execution subject of each process will not be described again in the following description of the gesture control method. Specifically, the specific process of the gesture control method is as follows:

[0097] S601: Obtain the first dynamic gesture video to be processed.

[0098] In this application, the first dynamic gesture video can be a user dynamic gesture video corresponding to the user performing a gesture operation. The first dynamic gesture video can be a video stream.

[0099] In some embodiments, if the electronic device is a mobile phone or other terminal device, the camera can send the captured first dynamic gesture video to the processor or other modules of the electronic device, enabling the processor to execute subsequent steps S602-S604, thereby realizing the gesture control function corresponding to the first dynamic gesture video. Alternatively, if the electronic device is a smart terminal corresponding to a smart home, the camera can send the captured first dynamic gesture video to the processing device of the smart terminal, enabling the processing device to execute subsequent steps S602-S604, thereby realizing the gesture control function corresponding to the first dynamic gesture video. Or, if the electronic device is a vehicle infotainment system, the camera in the vehicle can send the captured first dynamic gesture video to the vehicle infotainment system, enabling the system to execute subsequent steps S602-S604, thereby realizing the gesture control function corresponding to the first dynamic gesture video.

[0100] S602: Determine the first gesture feature based on the position of each hand key point in each video frame of the first dynamic gesture video.

[0101] In this application, the first gesture feature can be the gesture feature corresponding to the user gesture operation mentioned in this application. The type of the first gesture feature can be at least one of feature vector, matrix, and feature map; the key points of the hand can be wrist joint, fingertip, finger joint, etc., which are not limited in this application.

[0102] Wherein, when the type of the first gesture feature is a feature vector, the electronic device can determine the first gesture feature vector corresponding to the user's gesture operation based on the position of each hand key point in each video frame of the first dynamic gesture video. The first gesture feature vector may include multiple feature parts such as displacement feature part, static posture feature part, and hand orientation feature part. Specifically, the acquisition method of each feature part can be referred to as follows: (1) Displacement feature part, used to indicate the displacement of each hand key point in the first dynamic gesture video, for example, may include displacement information of n hand key points. (2) Static posture feature part, used to indicate hand posture, which may include the number of extended fingers and / or the shape fitted by the hand, etc. In some embodiments, the electronic device can determine the static posture feature data group corresponding to each video frame based on the position of each hand key point in each video frame, obtaining multiple static posture feature data groups; then, the average value between the corresponding feature data in the multiple static posture feature data groups is obtained, thereby obtaining each feature data in the static posture feature part. For example, in the first video frame, after the electronic device determines the coordinates of each key point of the hand, if the static posture feature data group (e.g., L1', L2') in the first video frame, the static posture feature data group (L1'', L2'') in the second video frame, and the static posture feature data group (L1''', L2'') in the third video frame are represented by the distance between multiple specific key points, the electronic device can calculate the average value of the corresponding static posture feature data in multiple static posture feature data groups, thereby obtaining the static posture feature part (L1, L2) in the first gesture feature vector. (3) Hand orientation feature part, used to indicate the orientation of the palm. In some embodiments, the electronic device can determine the hand orientation feature data group corresponding to each video frame based on the position of each hand key point in each video frame, thus obtaining multiple hand orientation feature data groups; then, it can obtain the average value between the corresponding feature data in the multiple hand orientation feature data groups to obtain each feature data in the hand orientation feature part. For example, the electronic device can represent the hand orientation feature data group by the coordinates of multiple palm feature points. For example, the hand orientation feature data group in the first video frame is (p1', p2'), the hand orientation feature data group in the second video frame is (p1'', p2''), and the hand orientation feature data group in the third video frame is (p1'', p2''). The electronic device can calculate the average value of the corresponding hand orientation feature data in the multiple hand orientation feature data groups to obtain the hand orientation feature part (p1, p2) in the first gesture feature vector.

[0103] Thus, after the electronic device determines the feature data of multiple feature components such as the displacement feature part, the static gesture feature part, and the hand orientation feature part based on the positions of each hand key point in each video frame, it can determine the first gesture feature vector.

[0104] For another example, in some other embodiments, the electronic device can also process the first dynamic gesture video through a sliding window algorithm, so that the obtained first gesture feature vector includes n (n is a positive integer) subsequence feature vectors. Specifically, the electronic device can first obtain n video segments based on a sliding window with a preset length of x1 video frames and a sliding step of x2 (both x1 and x2 are positive integers) video frames, that is, the electronic device can use the video segment between the i-th video frame and the (i + x1 - 1)-th (i≥1 and i is a positive integer) video frame corresponding to the sliding window as the h-th (1≤k<n and h is a positive integer) video segment; next, the electronic device can move the sliding window by x2 video frames and use the video segment between the (i + x2)-th video frame and the (i + x2 + x1 - 1)-th video frame corresponding to the sliding window as the (h + 1)-th video segment. Thus, n video segments can be obtained.

[0105] For example, if the length x1 of the sliding window is 10 and the sliding step x2 is 2, the electronic device can use the video segment between the 1st frame and the 10th frame in the first dynamic gesture video as the first video segment. Then, the electronic device can control the sliding window to move a distance of 2 video frames and use the video segment between the 3rd frame and the 12th frame as the second video segment; next, the electronic device can control the sliding window to move a distance of 2 video frames again and use the video segment between the 5th frame and the 14th frame as the third video segment. Similarly, the electronic device can control the sliding window to move 2 video frames in sequence to obtain n video segments.

[0106] It can be understood that the initial frame of the first video segment can be the 1st video frame in the first dynamic gesture video or any video frame in the first dynamic gesture video. For example, the video segment between the 1st frame and the x1-th frame in the first dynamic gesture video can be used as the first video segment, and then the sliding window is controlled to move x2 video frames in sequence to obtain n video segments; or the video segment between the i-th frame and the (i + x1 - 1)-th frame in the first dynamic gesture video can be used as the first video segment, and then the sliding window is controlled to move x2 video frames in sequence to obtain n video segments. This application does not make any limitations. Among them, the length x1 and the sliding step x2 of the sliding window can also be set arbitrarily, and this application does not make any limitations. Next, after the electronic device obtains n video segments based on the sliding window algorithm, it can determine the feature vectors (which can be described as subsequence feature vectors) corresponding to each video segment in the n video segments based on the positions of each hand key point in each video frame of each video segment.

[0107] Therefore, compared to directly classifying user gesture operations based on the first gesture feature vector corresponding to the first dynamic gesture video, this application classifies user gesture operations based on sub-sequence feature vectors corresponding to multiple video segments, allowing for similarity matching of gesture operations in each video segment. For example, if the sub-sequence feature vectors corresponding to gesture operations in multiple video segments all satisfy the similarity condition with the corresponding template feature vector in gesture operation A, then the complete user gesture operation type can be determined to be gesture operation A. Similarly, if only one or a very small number of video segments have sub-sequence feature vectors corresponding to gesture operations that satisfy the similarity condition with the corresponding template feature vector in gesture operation A, and multiple consecutive video segments have sub-sequence feature vectors corresponding to gesture operations that satisfy the similarity condition with the corresponding template feature vector in gesture operation B, then the complete user gesture operation type can be determined to be gesture operation B. It is evident that classifying user gesture operations based on sub-sequence feature vectors corresponding to multiple video segments can result in more accurate classification of user gesture operations.

[0108] S603: Determine the similarity between the first gesture feature and the preset template gesture features.

[0109] In this application, the types of the preset template gesture features are the same as the type of the first gesture feature.

[0110] Where both the first gesture feature and each template gesture feature are feature vectors, a template feature vector library can be pre-set in the electronic device, storing multiple template feature vectors and the corresponding control operations. Thus, when the first gesture feature vector corresponding to the user's gesture operation is determined through S602, the electronic device can determine the similarity between the first gesture feature vector and each template feature vector in the pre-set template feature vector library. The template feature vectors and their corresponding control operations can be stored before the electronic device leaves the factory, or they can be stored by the user through a custom mode; this application does not limit this.

[0111] In some embodiments, if the electronic device directly determines the complete first gesture feature vector based on the position of each hand key point in each video frame of the first dynamic gesture video, the electronic device can directly compare the first gesture feature vector with each template feature vector, for example, by directly calculating the Euclidean distance between the first gesture feature vector and each template feature vector.

[0112] However, in other embodiments, if the electronic device obtains n video segments of the first dynamic gesture video based on the sliding window algorithm, and obtains the subsequence feature vectors corresponding to each video segment in the n video segments respectively (that is, the first gesture feature vector includes n subsequence feature vectors), then the electronic device needs to calculate the similarity between each subsequence feature vector in the n subsequence feature vectors and each template feature vector (which can be described as subsequence similarity), so as to obtain the n subsequence similarities between the first gesture feature vector and each template feature vector.

[0113] S604: When the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition, the first control operation corresponding to the first template gesture feature is executed.

[0114] In this application, the first template gesture feature can be a specific template gesture feature that satisfies similarity conditions to the first gesture feature; the first control operation can be the control operation corresponding to the specific template gesture feature.

[0115] When both the first gesture feature and each template gesture feature are feature vectors, the electronic device can execute the first control operation corresponding to the first template feature vector if the similarity between the first gesture feature vector and the first template feature vector meets the similarity condition.

[0116] Specifically, in some embodiments, if the electronic device directly determines the complete first gesture feature vector based on the position of each hand key point in each video frame of the first dynamic gesture video, the electronic device can directly compare the first gesture feature vector with each template feature vector, and then use the specific template feature vector that meets the similarity condition as the first template feature vector.

[0117] In other embodiments, if the first gesture feature vector includes n sub-sequence feature vectors, and the electronic device further calculates the similarity of the first gesture feature vector with each template feature vector among the n sub-sequences, if the electronic device determines that k sub-sequence similarities among the n sub-sequences of the similarity between the first gesture feature vector and the first template feature vector satisfy the similarity condition, then the electronic device can determine that the similarity between the first gesture feature vector and the first template feature vector satisfies the similarity condition, thereby executing the first control operation corresponding to the first template feature vector. Here, k ≤ n, and k is a positive integer greater than a preset threshold number.

[0118] In this way, electronic devices can classify gesture operations in each video segment. For example, if multiple video segments all contain gesture operation A, it can be determined that the first gesture feature vector corresponding to that user gesture operation satisfies the similarity condition with the first template feature vector corresponding to gesture operation A. Similarly, if only one or a very small number of video segments contain gesture operation A, while multiple consecutive video segments contain gesture operation B, it can be determined that the first gesture feature vector corresponding to that user gesture operation satisfies the similarity condition with the first template feature vector corresponding to gesture operation B. Therefore, classifying user gesture operations based on the sub-sequence feature vectors corresponding to multiple video segments can make the classification results more accurate.

[0119] In some embodiments, after determining that the similarity between the first gesture feature and the first template gesture feature meets the similarity condition, the electronic device can send an instruction corresponding to the first control operation to the execution device corresponding to the first control operation, so that the execution device can respond to the instruction and execute the first control operation. For example, if the execution device is an in-vehicle device, the electronic device (e.g., a vehicle infotainment system) can send the instruction corresponding to the first control operation to the in-vehicle device, so that the in-vehicle device executes the first control operation; as another example, if the execution device is a smart home device, the electronic device (e.g., a smart terminal) can send the instruction corresponding to the first control operation to the smart home device, so that the smart home device executes the first control operation.

[0120] Thus, the above gesture control method can improve the accuracy of user gesture operation classification, thereby accurately realizing gesture control function. At the same time, the above method can also effectively reduce the consumption of computing and storage resources of electronic devices.

[0121] The following is combined with Figure 7 The flowchart shown provides a detailed explanation of the gesture control method provided in this application. It can be understood that in the embodiments of this application, Figure 7 The execution subject of each process in the illustrated method can be an electronic device, such as the aforementioned mobile phone 301, smart terminal, or in-vehicle system, etc. The execution subject of each process will not be described again in the following description of the gesture control method. Specifically, the specific process of the gesture control method is as follows:

[0122] S701: Start.

[0123] This application does not limit the start time of the provided gesture control method. For example, the function corresponding to the gesture control method provided in this application can be the default function of the electronic device, that is, the gesture control method provided in this application can start running after the electronic device is started; or, the function corresponding to the gesture control method provided in this application can be an optional function of the electronic device, that is, the electronic device can start executing the gesture control method provided in this application after detecting that the user has enabled the gesture control function. This application does not limit this.

[0124] S702: Acquire dynamic gesture video recorded by the camera.

[0125] In this application, the dynamic gesture video can be the complete first dynamic gesture video; or it can be a video segment obtained based on the sliding window algorithm. That is to say, the gesture control method provided in this application embodiment can either process the entire first dynamic gesture video directly, or process each video segment in the n video segments divided from the first dynamic gesture video separately.

[0126] S703: Identify key points of each hand and obtain the time series corresponding to each key point of the hand.

[0127] In some embodiments, the electronic device can identify key points of the user's hand in each video frame of a dynamic gesture video (which can be the complete first dynamic gesture video or a video segment divided from the first dynamic gesture video), and then obtain the time sequence corresponding to each key point based on the position of the key points in each video frame. For example, if the key point of the hand includes finger joint 1, and the position coordinate of finger joint 1 in the first video frame is p1, the position coordinate of finger joint 1 in the second video frame is p2, and the position coordinate of finger joint 1 in the third video frame is p3, then the time sequence corresponding to finger joint 1 is (p1, p2, p3). Similarly, the corresponding time sequence can also be determined for other key points (such as finger joint 2, wrist joint, etc.), which will not be elaborated here.

[0128] S704: Feature extraction is performed based on the time series corresponding to each key hand point to obtain the feature vector to be processed.

[0129] In this embodiment of the application, if the dynamic gesture video obtained in S702 is a complete first dynamic gesture video, then the feature vector to be processed can be the first gesture feature vector mentioned above; if the dynamic gesture video obtained in S702 is a video segment divided from the first dynamic gesture video, then the feature vector to be processed can be the subsequence feature vector corresponding to the video segment.

[0130] In some embodiments, feature extraction can be performed after determining the time series of each key hand point. For details, please refer to S602 above; further elaboration is not required here.

[0131] S705: Compare the feature vector to be processed with the preset feature vectors of each template.

[0132] In some embodiments, the electronic device may calculate the similarity between the feature vector of the feature to be processed and the feature vectors of each template in a variety of ways, such as Euclidean distance, cosine similarity, Manhattan distance, etc., which are not limited in this application.

[0133] In some embodiments, if the feature vector to be processed is the first gesture feature vector corresponding to the complete first dynamic gesture video, the electronic device can calculate the similarity between the first gesture feature vector and each template feature vector.

[0134] In other embodiments, if the feature vector to be processed is a subsequence feature vector corresponding to a video segment, the electronic device needs to calculate the similarity between each template feature vector and the subsequence feature vector.

[0135] S706: Determine if there exists a preset template feature vector and a feature vector to be processed whose Euclidean distance is less than a preset distance threshold. If yes, proceed to S707; otherwise, proceed to S709.

[0136] In some embodiments, if the electronic device uses Euclidean distance to represent the similarity between the feature vector of the feature to be processed and the feature vectors of each template, the similarity condition may include the Euclidean distance being less than a preset distance threshold. Alternatively, the similarity condition may also be that the Euclidean distance is less than a preset distance threshold, and the Euclidean distance is the smallest Euclidean distance among all Euclidean distances.

[0137] S707: Classify the feature vector to be processed into the preset first template feature vector with the smallest Euclidean distance.

[0138] In some embodiments, if the feature vector to be processed is the first gesture feature vector corresponding to the complete first dynamic gesture video, then S708 can be executed directly after the first template feature vector is determined.

[0139] In other embodiments, if the feature vector to be processed is a sub-sequence feature vector corresponding to a video segment, the electronic device also needs to determine that the sub-sequence feature vectors corresponding to other video segments are all first template feature vectors, or that the number of sub-sequence feature vectors belonging to the first template feature vector is greater than a preset number threshold before S708 can be executed.

[0140] S708: Execute the first control operation corresponding to the first template feature vector.

[0141] In some embodiments, an electronic device may control an execution device corresponding to a first control operation to perform the first control operation.

[0142] S709: End.

[0143] In some embodiments, after the electronic device performs a first control operation, it indicates that a gesture control method has been completed.

[0144] Thus, the above gesture control method can improve the accuracy of user gesture operation classification, thereby accurately realizing gesture control function. At the same time, the above method can also effectively reduce the consumption of computing and storage resources of electronic devices.

[0145] For example, refer to Figure 8 As shown, after receiving a user's dynamic gesture video, the vehicle's infotainment system in vehicle 501 can perform hand keypoint recognition on multiple video frames in the video, thereby obtaining the time series corresponding to each hand keypoint. Next, the system can extract gesture features from the time series corresponding to each hand keypoint, thus obtaining a feature vector to be processed. For example, Figure 8 The feature vector to be processed shown can be composed of three parts: static posture features (r1, r2...), hand displacement features (s1, s2...), and hand orientation features (u1, u2...). Finally, the vehicle's infotainment system can calculate the Euclidean distance between the feature vector to be processed and each template feature vector in the preset template feature vector library. The preset template feature vector library stores each gesture operation, the corresponding template feature vector, and the corresponding control operation. When the vehicle's infotainment system determines that the similarity between the feature vector to be processed and a certain first template feature vector meets the similarity condition, the infotainment system can determine that the feature vector to be processed belongs to that first template feature vector. Wherein, if... Figure 8 When the feature vector to be processed shown is the first gesture feature vector, and the first gesture feature vector belongs to the first template feature vector, the vehicle system can directly control the in-vehicle equipment to perform the corresponding first control operation. Or, if Figure 8 The feature vector to be processed shown is a sub-sequence feature vector. When the vehicle system determines that all other sub-sequence feature vectors belong to the first template feature vector, it can control the in-vehicle devices to perform the corresponding first control operation. In this way, while realizing gesture control operations on in-vehicle devices, it can also effectively reduce the consumption of computing and storage resources of the vehicle system.

[0146] Furthermore, in some embodiments, the electronic device may also include a custom mode. When the electronic device is in custom mode, upon acquiring a second dynamic gesture video recorded by the user, it can determine the second gesture feature vector corresponding to the second dynamic gesture video based on the positions of each hand key point in each video frame. This second gesture feature vector can then be added to a preset template feature vector library. Next, the user can choose to bind a new second control operation to the second gesture feature vector, or replace other template feature vectors corresponding to the second control operation with the second gesture feature vector. Thus, when the electronic device receives an operation from the user associating the second gesture feature vector with a second control operation, it can respond to the user's editing operation by storing the association between the second gesture feature vector and the second control operation, thereby making the template feature vector corresponding to the second control operation the second gesture feature vector.

[0147] Specifically. The following will combine... Figure 9 The flowchart shown provides a detailed explanation of the gesture control method for an electronic device in custom mode. It can be understood that, in the embodiments of this application, Figure 9 The execution subject of each process in the method shown can be an electronic device, such as the aforementioned mobile phone 301, smart terminal, or in-vehicle system, etc. The execution subject of each process will not be described again in the following description of the gesture control method. Specifically, the specific process of the gesture control method is as follows:

[0148] S901: Acquire multiple third-party dynamic gesture videos recorded by the camera.

[0149] In some embodiments, when the electronic device is in a custom mode, it can acquire multiple third-party dynamic gesture videos recorded by the user.

[0150] S902: Identify key points of each hand and obtain the time series corresponding to each key point of the hand.

[0151] In each third dynamic gesture video, the electronic device can acquire the time sequence corresponding to each key point of the hand, as described in S703 above, and will not be repeated here.

[0152] S903: Based on the time series corresponding to each gesture key point, feature extraction is performed to obtain multiple third gesture feature vectors.

[0153] In some embodiments, for each third dynamic gesture video, the electronic device can determine the third gesture feature vector corresponding to each third dynamic gesture video based on the position of each gesture key point in multiple video frames. For details, please refer to S602 above, which will not be repeated here.

[0154] S904: Determine whether the number of video recordings has reached the maximum threshold; if yes, proceed to S905; otherwise, proceed to S901.

[0155] In some embodiments, if the number of video recordings reaches the maximum threshold, it indicates that the user has completed the video recording input, and the process can proceed to S905; conversely, if the number of video recordings does not reach the maximum threshold, the user can be reminded to continue recording the video, and the process can proceed to S901.

[0156] In some embodiments, the electronic device may not determine whether the number of video recordings has reached the maximum threshold, but instead directly execute S905 when it detects that the user has ended recording.

[0157] S905: Add the average feature vector of each third gesture feature vector corresponding to each third dynamic gesture video to the preset template feature vector library.

[0158] In some embodiments, after determining multiple third gesture feature vectors, the electronic device can store the average feature vector of the multiple third gesture feature vectors in a preset template feature vector library.

[0159] S906: Associate the average eigenvector with the corresponding control operation.

[0160] In some embodiments, the user can choose to bind a new third control operation to the average feature vector, or replace other template feature vectors corresponding to the third control operation with the average feature vector. Thus, when the electronic device receives a user's operation to associate the average feature vector with a third control operation, the electronic device can respond to the user's editing operation by storing the association between the average feature vector and the third control operation, thereby making the template feature vector corresponding to the third control operation the average feature vector.

[0161] Thus, the above method also allows users to maintain their own preset template feature vector library, thereby achieving scalability of the template feature vector library. For example, refer to... Figure 10 As shown, when the vehicle's infotainment system is in custom mode, the vehicle 501 can determine the third gesture feature vector corresponding to each of the multiple third dynamic gesture videos after receiving them, thereby obtaining the average feature vector. For example, Figure 10The average feature vector shown can be composed of three parts: static posture features (r1, r2...), hand displacement features (s1, s2...), and hand orientation features (u1, u2...). Next, the vehicle's infotainment system can update the preset template feature vector library based on this average feature vector. For example, the system can receive a user's editing operation, binding the average feature vector to a user-defined gesture and a new control operation; or, the system can receive a user's editing operation, binding the average feature vector to a user-defined gesture and an existing control operation, and unbinding the bound existing control operation from the original template feature vector. This allows users to maintain their own preset template feature vector library and customize gesture control functions.

[0162] It is understood that the gesture control method provided in this application can be applied to any scenario where gesture-based air control can be achieved. Specifically, when the gesture control method provided in this application is used to control in-vehicle equipment, the vehicle may include a camera, a vehicle infotainment system, and an execution device. See reference [link to relevant documentation]. Figure 11 As shown, vehicle 501 may include camera 5011, vehicle infotainment system 5012, and execution device 5013. Camera 5011 can capture dynamic gesture video of the user and send it to vehicle infotainment system 5012. Vehicle infotainment system 5012 can determine a first gesture feature based on the position of each hand key point in each video frame of the dynamic gesture video. Furthermore, vehicle infotainment system 5012 can calculate the similarity between the first gesture feature and preset template gesture features, and if the similarity between the first gesture feature and the preset first template gesture features meets the similarity condition, control the execution device 5013 in vehicle 501 corresponding to the first template gesture feature to perform a corresponding first control operation. Execution device 5013 is used to execute the first control operation in response to the control command from vehicle infotainment system 5012. Thus, through the interaction between the various modules in vehicle 501, vehicle 501 can implement the gesture control method mentioned in this application.

[0163] Specifically, the functions of the vehicle infotainment system 5012 are as follows:

[0164] 5012a: Running gesture key point recognition model.

[0165] Among them, the gesture key point recognition model can be used to identify the key points of the hand in each video frame, thereby outputting the time series corresponding to each key point of the hand.

[0166] 5012b: Run the gesture feature vector extraction algorithm.

[0167] Among them, the gesture feature vector extraction algorithm can be used to determine the first gesture feature vector based on the time series corresponding to each key point of the hand.

[0168] 5012c: Determine if in custom mode? If yes, proceed to 5012e; otherwise, proceed to 5012d.

[0169] When the vehicle infotainment system 5012 is in custom mode, it allows users to change the template feature vector library; when the vehicle infotainment system 5012 is not in custom mode, it can classify the user's gesture operations.

[0170] 5012d: Run the dynamic gesture classification algorithm.

[0171] Among them, the dynamic gesture classification algorithm can be used to calculate the similarity between the first gesture feature vector and the feature vectors of each template, thereby classifying the user's first gesture feature vector.

[0172] The specific implementation process of 5012a to 5012d can be referred to the above. Figure 7 As shown, it will not be elaborated further here.

[0173] 5012e: Change the template feature vector library.

[0174] When the vehicle system 5012 is in custom mode, it can add the first gesture feature vector obtained by 5012b to the preset template feature vector library.

[0175] In this way, through the above process, the vehicle system 5012 can recognize the user's gesture operation, thereby controlling the execution device 5013 to perform the corresponding control operation or customize and update the template feature vector library.

[0176] Furthermore, in some embodiments, this application also provides a computer program product, wherein the computer program product includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device causes the electronic device to perform the gesture control method mentioned in this application.

[0177] In other embodiments, this application also provides a storage medium in which instructions are stored on the readable storage medium, and when executed on an electronic device, the instructions cause the electronic device to perform the gesture control method mentioned in this application.

[0178] In other embodiments, this application also provides an electronic device including a memory and a processor, wherein the memory is coupled to the processor. The memory stores computer program code / instructions, and when the computer program code / instructions are executed by the processor, the electronic device can perform the gesture control method mentioned in this application.

[0179] The following is combined Figure 12 The diagram illustrates the hardware structure of an electronic device 1200 according to an embodiment of this application. Figure 12 As shown, the electronic device 1200 may include one or more processors 1202, system control logic 1201 connected to at least one of the processors 1202, system memory 1205 connected to the system control logic 1201, memory 1203 connected to the system control logic 1201, and network interface 1207 connected to the system control logic 1201.

[0180] It is understood that the structures illustrated in the embodiments of this application do not constitute a limitation on the only possible implementation of the electronic device 1200. In other embodiments of this application, the electronic device 1200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0181] Processor 1202 may include one or more single-core or multi-core processors. In some embodiments, processor 1202 may include any combination of general-purpose processors and special-purpose processors (e.g., application processors, baseband processors, etc.). It is understood that in this embodiment, processor 1202 may be configured to execute executable instructions 1204 stored in memory 1203 to implement the gesture control method of this embodiment. When at least one instruction is executed in processor 1202, electronic device 1200 implements the gesture control method of this embodiment.

[0182] System control logic 1201 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1202 and / or any suitable device or component communicating with system control logic 1201. System control logic 1201 may include one or more memory controllers to provide an interface to system memory 1205. System memory 1205 may be used to load and store data and / or instructions. In some embodiments, system memory 1205 of electronic device 1200 may include any suitable volatile memory, such as suitable dynamic random access memory.

[0183] The memory 1203 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the memory 1203 may include any suitable volatile memory and / or any suitable non-volatile storage device, such as a random access memory (RAM) and / or a cache memory, and may further include a read-only memory (ROM).

[0184] The memory 1203 may include a portion of the storage resources on the device on which the electronic device 1200 is installed, or it may be accessible by the device, but is not necessarily part of the device. For example, the memory 1203 may be accessed over a network via a network interface 1207.

[0185] Specifically, system memory 1205 and storage 1203 may each include a temporary copy and a permanent copy of instruction 1204. Instruction 1204 may include, when executed by at least one of processors 1202, causing electronic device 1200 to implement the gesture control method of the embodiments of this application. In some embodiments, instruction 1204, hardware, firmware, and / or its software components may additionally / alternatively be placed in system control logic 1201, network interface 120, and / or processor 1202.

[0186] Network interface 1207 may include a transceiver for providing a radio interface to electronic device 1200, thereby enabling communication with any other suitable device (such as a front-end module, antenna, etc.) via one or more networks. In some embodiments, network interface 1207 may be integrated into other components of electronic device 1200. For example, network interface 1207 may be integrated into at least one of processor 1202, system memory 1205, memory 1203, and firmware device (not shown) with instructions.

[0187] The network interface 1207 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1207 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.

[0188] The electronic device 1200 may further include an input / output (I / O) device 1206. The I / O device 1206 may include a user interface enabling a user to interact with the electronic device 1200; the peripheral component interface is designed to allow peripheral components to also interact with the electronic device 1200. In some embodiments, the electronic device 1200 may also include sensors for determining at least one of environmental conditions and location information related to the electronic device 1200.

[0189] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.

[0190] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0191] In some embodiments, the sensor may include, but is not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 1207 to communicate with components of the positioning network (e.g., global positioning system (GPS) satellites).

[0192] The embodiments disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0193] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.

[0194] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0195] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, optical discs, magneto-optical disks, read-only memory (ROM), random access memory (RAM), magnetic cards or optical cards, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other propagation signals. Therefore, machine-readable media include any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0196] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0197] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0198] It should be noted that in the examples and description of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0199] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of this application.

Claims

1. A gesture control method applied to electronic devices, characterized in that, The method includes: Obtain the first dynamic gesture video to be processed; The first gesture feature is determined based on the position of each key hand point in each video frame of the first dynamic gesture video; The similarity between the first gesture feature and the preset template gesture features is determined. If the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition, the first control operation corresponding to the first template gesture feature is executed.

2. The method according to claim 1, characterized in that, The type of the first gesture feature includes at least one of feature vector, matrix, and feature map; and, The type of the first gesture feature is the same as the type of each template gesture feature.

3. The method according to claim 2, characterized in that, The type corresponding to the first gesture feature and the gesture features of each template is a feature vector. The step of determining the first gesture feature based on the position of each hand key point in each video frame of the first dynamic gesture video includes: Based on the position of each key hand point in each video frame of the first dynamic gesture video, the first gesture feature vector is determined. Determining the similarity between the first gesture feature and the preset template gesture features includes: The similarity between the first gesture feature vector and each template feature vector in the preset template feature vector library is determined; and... If the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition, the first control operation corresponding to the first template gesture feature is executed, including: If the similarity between the first gesture feature vector and the first template feature vector meets the similarity condition, the first control operation corresponding to the first template feature vector is executed.

4. The method according to claim 3, characterized in that, The feature vectors of all templates in the preset template feature vector library have the same dimension; and... The dimension of the first gesture feature vector is the same as the dimension of each template feature vector.

5. The method according to claim 4, characterized in that, The methods for obtaining each feature data in the first gesture feature vector include: Based on the position of each hand key point in each video frame, the gesture feature data group corresponding to each video frame is determined, resulting in multiple gesture feature data groups; The average value of corresponding feature data in the plurality of gesture feature data groups is obtained to obtain each feature data in the first gesture feature vector.

6. The method according to claim 4, characterized in that, The feature components of the first gesture feature vector include at least one of a static posture feature component, a displacement feature component, and a hand orientation feature component; wherein, The static posture feature portion is used to indicate hand posture, which includes at least one of the following: the number of fingers extended from the hand and the shape fitted to the hand. The displacement feature portion is used to indicate the displacement of each key hand point in the first dynamic gesture video; The hand orientation feature is used to indicate the orientation of the palm.

7. The method according to claim 3 or 4, characterized in that, The first gesture feature vector includes n sub-sequence feature vectors; and, The step of determining the first gesture feature vector based on the position of each hand key point in each video frame of the first dynamic gesture video further includes: Obtain n video segments based on a preset sliding window. The length of the sliding window is x1 video frames, and the sliding step of the sliding window is x2 video frames. Both x1 and x2 are positive integers; Based on the positions of each hand key point in each video frame of each video segment, respectively determine the subsequence feature vectors corresponding to each video segment in the n video segments.

8. The method according to claim 7, characterized in that, The obtaining of the n video segments based on the preset sliding window includes: Regarding the video segment between the i-th video frame and the (i + x1 - 1)-th video frame corresponding to the sliding window as the h-th video segment; Move the sliding window by x2 video frames, and regard the video segment between the (i + x2)-th video frame and the (i + x2 + x1 - 1)-th video frame corresponding to the sliding window as the (h + 1)-th video segment; where, i ≥ 1 and i is a positive integer, 1 ≤ h < n and h is a positive integer.

9. The method according to claim 8, characterized in that, The determining of the similarities between the first gesture feature vector and each template feature vector in the preset template feature vector library includes: Respectively calculate the subsequence similarities between each subsequence feature vector in the n subsequence feature vectors and each template feature vector, and obtain the n subsequence similarities between the first gesture feature vector and each template feature vector.

10. The method according to claim 9, characterized in that, When the similarity between the first gesture feature vector and the first template feature vector meets the similarity condition, execute the first control operation corresponding to the first template feature vector, including: When k of the n subsequence similarities between the first gesture feature vector and the first template feature vector meet the similarity condition, determine that the similarity between the first gesture feature vector and the first template feature vector meets the similarity condition, and execute the first control operation corresponding to the first template feature vector; where, k ≤ n and k is a positive integer greater than the preset quantity threshold.

11. The method according to claim 4, characterized in that, The determining of the similarities between the first gesture feature vector and each template feature vector in the preset template feature vector library includes: Respectively calculate the Euclidean distances between the first gesture feature vector and each template feature vector; Regard each Euclidean distance as the similarity between the first gesture feature vector and the corresponding template feature vector; and, The similarity between the first gesture feature vector and the first template feature vector meeting the similarity condition includes: The Euclidean distance between the first gesture feature vector and the first template feature vector is less than the preset distance threshold.

12. The method according to claim 4, characterized in that, The electronic device includes a custom mode, and the method further includes: When the electronic device is in the custom mode, obtain a second dynamic gesture video; Based on the positions of each hand key point in each video frame of the second dynamic gesture video, determine the second gesture feature vector corresponding to the second dynamic gesture video; Add the second gesture feature vector to the preset template feature vector library; Obtain the operation of the user associating the second gesture feature vector with a second control operation, and store the association relationship between the second gesture feature vector and the second control operation.

13. The method according to claim 12, characterized in that, The method further includes: When the electronic device is in the custom mode, multiple third dynamic gesture videos are acquired; Based on the position of each hand key point in each video frame of each third dynamic gesture video, the feature vector of each third gesture corresponding to each third dynamic gesture video is determined. Add the average feature vector between the feature vectors of each third gesture corresponding to each third dynamic gesture video to the preset template feature vector library; The user's operation of associating the average feature vector with the third control operation is obtained, and the association relationship between the average feature vector and the third control operation is stored.

14. The method according to claim 1, characterized in that, The step of executing the first control operation corresponding to the first template gesture feature when the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition further includes: When the similarity between the first gesture feature and the first template gesture feature meets the similarity condition, the electronic device sends the instruction corresponding to the first control operation to the execution device corresponding to the first control operation; wherein... The instruction corresponding to the first control operation is used to cause the execution device to perform the first control operation.

15. A vehicle, characterized in that, include: A camera used to capture videos of dynamic gestures; The vehicle's infotainment system is used to determine the first gesture feature based on the position of each hand key point in each video frame of the dynamic gesture video; and, The vehicle system is also used to calculate the similarity between the first gesture feature and the preset template gesture features, and If it is determined that the similarity between the first gesture feature and the preset first template gesture feature meets the similarity condition, the vehicle system is further configured to control the execution device in the vehicle corresponding to the first template gesture feature to perform a corresponding first control operation. An execution device is used to execute the first control operation in response to the control command of the vehicle system.

16. An electronic device, characterized in that, include: A memory and a processor, the memory being coupled to the processor; the memory being used to store computer program code / instructions; when the computer program code / instructions are executed by the processor, causing the electronic device to perform the gesture control method according to any one of claims 1 to 14.

17. A readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the gesture control method according to any one of claims 1 to 14.

18. A computer program product, characterized in that, include: Computer instructions, when executed on an electronic device, cause the electronic device to perform the gesture control method according to any one of claims 1 to 14.