Machine vision-based gesture recognition method and system for intelligent fan
By incorporating a machine vision module and deep learning model into the smart fan, accurate recognition and detailed analysis of user gestures are achieved, solving the problem of insufficient gesture recognition accuracy in existing technologies and improving user experience and intelligent control of the fan.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SHENZHEN ZHONGZHI LIANCHENG TECHNOLOGY CO LTD
- Filing Date
- 2024-11-11
- Publication Date
- 2026-05-15
AI Technical Summary
Existing gesture recognition methods cannot accurately analyze changes in the waving rate and amplitude between similar gestures, resulting in decreased gesture recognition accuracy and affecting the user experience.
The smart fan uses a built-in machine vision module to acquire real-time images of user gestures, performs gesture background cropping and temporal feature point extraction, uses a deep learning model to recognize gestures, and combines image processing and machine learning algorithms to perform vector transformation and difference refinement analysis of key feature points at the gesture level, generating corresponding operation control commands.
It achieves accurate recognition of user gestures, capable of recognizing both slight and rapid waving gestures, thus enhancing the user experience. The fan can automatically adjust its speed and direction based on the user's gestures, providing a flexible control method.
Smart Images

Figure CN2024131171_15052026_PF_FP_ABST
Abstract
Description
A machine vision-based intelligent fan gesture recognition method and system Technical Field
[0001] This invention relates to the field of gesture recognition technology, and in particular to a smart fan gesture recognition method and system based on machine vision. Background Technology
[0002] Gesture recognition technology has garnered widespread attention due to its convenience and contactless nature, particularly in home appliances such as fans. Gesture-based control methods can enhance user experience and improve the intelligence of devices. Machine vision-based gesture recognition technology uses cameras to capture user gestures in real time and analyzes them through image processing and pattern recognition algorithms to achieve intelligent control of fans. This approach not only avoids the need for physical contact, improving flexibility and comfort, but also enables multiple functions such as adjusting fan speed and switching modes through simple gestures. However, existing gesture recognition methods cannot accurately analyze changes in the waving speed and amplitude between similar gestures, leading to decreased gesture recognition accuracy and impacting the user experience.
[0003] Summary of the Invention
[0004] Therefore, the present invention needs to provide a machine vision-based intelligent fan gesture recognition method and system to solve at least one of the above-mentioned technical problems.
[0005] To achieve the above objectives, a machine vision-based intelligent fan gesture recognition method includes the following steps:
[0006] Step S1: Obtain a set of real-time user gesture images through the machine vision module built into the smart fan, and perform gesture background cropping on the set of real-time user gesture images to obtain a set of user gesture background cropped images; perform gesture recognition and temporal feature point extraction on the set of user gesture background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture, wherein the user gestures include up, down and rotation gestures.
[0007] Step S2: Perform gesture-level key feature point vector transformation on the gesture part temporal change feature point set corresponding to each user's gesture action to obtain the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action; based on the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action, perform similar gesture action pair filtering and division on the gesture part temporal change feature point set corresponding to each similar gesture action to obtain the gesture part temporal feature point set corresponding to each similar gesture action; perform similar gesture difference refinement analysis on the gesture part temporal feature point set corresponding to each similar gesture action to obtain the smart fan user gesture difference refinement recognition result, which includes slightly upward, slightly downward, slightly rotating waving gesture actions as well as fast upward, fast downward, fast rotating waving gesture actions;
[0008] Step S3: Perform gesture operation control response for each user gesture in the refined recognition result of the smart fan user gesture differences, so as to generate the smart fan operation control command corresponding to each user gesture.
[0009] Step S4: Apply the smart fan operation control command corresponding to each user's specific gesture to the smart fan's built-in control system to execute the corresponding smart fan gesture action control working mode.
[0010] Furthermore, the present invention also provides a machine vision-based intelligent fan gesture recognition system for executing the machine vision-based intelligent fan gesture recognition method described above. The machine vision-based intelligent fan gesture recognition system includes:
[0011] The gesture action temporal feature point extraction module is used to acquire a set of real-time user gesture action images through the machine vision module built into the smart fan, and to perform gesture background cropping processing on the set of real-time user gesture action images to obtain a set of user gesture action background cropped images; and to perform gesture action recognition and temporal feature point extraction on the set of user gesture action background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture action, wherein the user gesture actions include upward, downward and rotation gesture actions;
[0012] The similar gesture action difference refinement and recognition module is used to transform the gesture-level key feature point vector of the gesture part temporal change feature point set corresponding to each user's gesture action to obtain the gesture part hierarchical representation key feature point vector of each user's gesture action; based on the gesture part hierarchical representation key feature point vector of each user's gesture action, similar gesture action pairs are filtered and divided into similar gesture action pairs to obtain the gesture part temporal feature point set corresponding to each similar gesture action; based on the gesture part temporal feature point set corresponding to each similar gesture action, similar gesture difference refinement analysis is performed to obtain the smart fan user gesture difference refinement and recognition results, which include slightly upward, slightly downward, slightly rotating waving gesture actions and quickly upward, quickly downward, quickly rotating waving gesture actions.
[0013] The user gesture refinement action control response module is used to perform gesture action operation control response for each specific action of the user gesture within the refined recognition result of the user gesture difference of the smart fan, so as to generate the smart fan operation control command corresponding to each specific action of the user gesture.
[0014] The user gesture control command response and execution module is used to apply the smart fan operation control command corresponding to each user gesture to the built-in control system of the smart fan, so as to execute the corresponding smart fan gesture action control working mode.
[0015] The beneficial effects of this invention are:
[0016] 1. The intelligent fan gesture recognition method based on machine vision proposed in this invention, compared with the prior art, has the advantage of capturing a set of real-time user gesture images using a built-in machine vision module. This allows for real-time monitoring of user actions and corresponding operations based on those gestures, enhancing the user experience and making fan use more intuitive and convenient. By performing gesture background cropping on each frame of the user's real-time gesture image set, interfering backgrounds can be removed, retaining only the gesture-related parts. This processing not only improves the accuracy of subsequent action recognition but also reduces the influence of irrelevant information on the algorithm, making gesture recognition more focused. This background cropping enables better recognition of user intentions under various environmental conditions. Simultaneously, by performing gesture recognition and temporal feature point extraction on a cropped image set of user gesture actions, a set of temporal change feature points for gesture parts is obtained. The key to this process is the ability to identify different gesture actions (such as upward, downward, and rotational movements) and accurately track their changes. This extraction of temporal features enables an understanding of the dynamic characteristics of gesture actions, thereby better predicting user needs and intentions. This allows for more flexible control methods in smart home scenarios, such as quickly adjusting fan speed, slightly adjusting fan speed, and switching fan modes, greatly improving user convenience and satisfaction. Secondly, by performing gesture-level key feature point vector transformation on the temporal change feature point set corresponding to each user's gesture actions, the abstract gesture key feature points in different hierarchical representations can be converted into a data vector format that can be processed by machine learning algorithms. This process effectively transforms the set of temporally changing gesture action feature points into a fixed-length vector, thus improving the ability of subsequent processing to distinguish different gesture actions. By filtering and classifying similar gesture pairs based on the key feature point vectors representing the hierarchical representation of the gesture parts corresponding to each user's gesture, the core of this process lies in classifying similar gestures through similarity analysis of feature vectors. This not only makes the processing of user input more intelligent, but also provides a clearer reference for subsequent gesture recognition, enabling the identification of subtle differences between gestures and thus better understanding of user intent.Furthermore, by performing a detailed analysis of the differences between similar gestures based on the temporal feature points of the corresponding gesture parts, the system can more accurately understand the user's intentions. This detailed analysis not only considers multiple dimensions such as the speed, amplitude, and direction of the gesture, but also identifies individual differences in how users execute gestures. This detailed recognition helps distinguish between actions such as "slight upward" and "rapid upward," ensuring that the smart fan can respond accordingly. This allows for adjustments to parameters such as fan speed and direction based on subtle user operations, and precise analysis of the differences in the waving speed and amplitude of various similar gestures, thereby improving the accuracy of recognizing gestures such as "slight upward" and "rapid upward." Then, by performing gesture operation control responses for each specific user gesture within the detailed recognition results of the smart fan's user gesture differences, this stage transforms the user's gestures into specific operation control commands, enabling the user to interact with the smart device in a more natural way. For example, the system can adjust the fan speed when a slight waving gesture is recognized, and execute a command to quickly adjust the fan speed when a rapid waving gesture is recognized. Finally, by applying the smart fan operation control command corresponding to each user gesture to the fan's built-in control system, real-time control of the smart fan's wind speed is achieved. The successful implementation of this process not only ensures the effective operation of the smart fan but also enables the device to flexibly adapt to different user needs. Through real-time control feedback, the smart fan can automatically adjust its operating mode to match the user's gesture commands, thereby achieving various functions such as slight or rapid speed adjustment and wind direction adjustment. This intelligent control method greatly enhances the user's reliance on and trust in the smart fan, enabling it to efficiently and more accurately identify and respond to corresponding differentiated gesture actions.
[0017] 2. The intelligent fan gesture recognition system based on machine vision proposed in this invention consists of a gesture action timing feature point extraction module, a similar gesture action difference refinement and recognition module, a user gesture refinement action control response module, and a user gesture control command response execution module. It can realize any intelligent fan gesture recognition method based on machine vision described in this invention. It is used to combine the operations between computer programs running on each module to realize the intelligent fan gesture recognition method based on machine vision. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower input, and can quickly and effectively provide a more accurate and efficient intelligent fan gesture recognition process based on machine vision, thereby simplifying the operation process of the intelligent fan gesture recognition system based on machine vision. Attached Figure Description
[0018] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0019] Figure 1 is a flowchart illustrating the steps of the intelligent fan gesture recognition method based on machine vision according to the present invention.
[0020] Figure 2 is a detailed flowchart of step S1 in Figure 1. Detailed Implementation
[0021] To achieve the above objectives, please refer to Figures 1 and 2. This invention provides a machine vision-based intelligent fan gesture recognition method, comprising the following steps:
[0022] Step S1: Obtain a set of real-time user gesture images through the machine vision module built into the smart fan, and perform gesture background cropping on the set of real-time user gesture images to obtain a set of user gesture background cropped images; perform gesture recognition and temporal feature point extraction on the set of user gesture background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture, wherein the user gestures include up, down and rotation gestures.
[0023] In this embodiment of the invention, a machine vision module built into the smart fan is used to capture real-time user gesture images. This module is equipped with a high-resolution camera and an image processing unit, capable of capturing corresponding user gesture images at a rate of 30 frames per second, thereby obtaining a set of real-time user gesture images. Edge contour planning and connection analysis is performed on each frame of the previously monitored real-time user gesture image set to further analyze the edge contours, thereby generating edge contour region lines corresponding to each frame of the gesture image. Cropping processing is then performed on the corresponding gesture images between the gesture region and the background region to determine the bounding rectangle region of the gesture using the contour region lines. Subsequently, the original image is cropped using this rectangular region. This operation employs an image cropping algorithm, retaining only the gesture portion and removing background information. After cropping, a set of cropped user gesture images with background is obtained. Then, by performing specific gesture recognition analysis and corresponding temporal change feature point extraction on the previously cropped user gesture action background image set, a deep learning model (such as a convolutional neural network) is used to extract features from the cropped images to identify gesture actions. The model is trained on a large number of gesture samples and can accurately identify gesture actions such as up, down, and rotation. Based on the previously identified specific gesture actions, temporal change feature point extraction is performed to obtain the corresponding gesture part temporal change feature point. The extracted feature points are arranged in chronological order to form a set of temporal feature point points, representing the dynamic changes of gesture actions. Finally, a set of temporal change feature point points corresponding to each user gesture action is obtained, where user gesture actions include up, down, and rotation gesture actions.
[0024] Step S2: Perform gesture-level key feature point vector transformation on the gesture part temporal change feature point set corresponding to each user's gesture action to obtain the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action; based on the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action, perform similar gesture action pair filtering and division on the gesture part temporal change feature point set corresponding to each similar gesture action to obtain the gesture part temporal feature point set corresponding to each similar gesture action; perform similar gesture difference refinement analysis on the gesture part temporal feature point set corresponding to each similar gesture action to obtain the smart fan user gesture difference refinement recognition result, which includes slightly upward, slightly downward, slightly rotating waving gesture actions as well as fast upward, fast downward, fast rotating waving gesture actions;
[0025] In this embodiment of the invention, the previously aggregated set of temporal change feature points corresponding to each user's gesture action is divided into different levels of representation within the gesture feature space. Image processing technology and deep learning algorithms are used to divide the gesture feature points into multiple levels based on the complexity and characteristics of the user's gesture action. The corresponding levels are divided into different parts such as fingers, palms, and forearms. By setting an importance threshold, feature points that significantly affect the recognition effect of the gesture action are selected. At the same time, feature point vectors are transformed. This process uses a vectorization method to transform the coordinate values and statistics of each key feature point into spatial vectors to form a unified feature vector, ensuring that all feature point vectors have the same scale. The coordinate values of each feature point are combined according to the defined feature dimensions (such as x, y, z coordinates, velocity, angle, etc.) to obtain the key feature point vectors representing the gesture parts of each user's gesture action. By combining the key feature point vectors of the gesture parts corresponding to each user's gesture action obtained from the previous transformation, and using similarity calculation algorithms (such as cosine similarity or Euclidean distance), the gesture vectors are compared to identify the gesture action sequence pairs with the same similarity. The feature point set corresponding to the gesture action in each group is then filtered and divided to obtain the temporal feature point set of the gesture parts corresponding to each similar gesture action, thus obtaining the temporal feature point set of the gesture parts corresponding to each similar gesture action. Then, by combining the temporal feature point sets of the gesture parts corresponding to the previously segmented similar gestures, a detailed analysis of the differences between similar gestures is performed. This process mainly analyzes the subtle differences between similar gestures. For example, for the two gestures of "slightly upward" and "quickly upward", the differences in the temporal changes of their key feature point vectors are compared. By conducting in-depth analysis of the waving rate and waving amplitude changes of the key feature point vectors corresponding to the gestures, the subtle differences in the gesture feature performance of the two types of gestures can be extracted, thereby identifying and dividing them to generate corresponding gesture difference recognition results, covering slightly upward, slightly downward, and slightly rotating waving gestures as well as quickly upward, quickly downward, and quickly rotating waving gestures, and finally obtaining the detailed recognition results of smart fan user gesture differences.
[0026] Step S3: Perform gesture operation control response for each user gesture in the refined recognition result of the smart fan user gesture differences, so as to generate the smart fan operation control command corresponding to each user gesture.
[0027] In this embodiment of the invention, the control response is refined based on the previously identified differences in user gestures of the smart fan. This refines the identification results to include specific actions of each user gesture (e.g., slight upward, slight downward, slight rotating gestures, and rapid upward, rapid downward, and rapid rotating gestures). In specific operations, a rule engine processes the identification results. For example, a slight upward gesture corresponds to "increase wind speed at low speed," a slight downward gesture corresponds to "decrease wind speed at low speed," a slight rotating gesture corresponds to "change wind direction at low speed," a rapid upward gesture corresponds to "increase wind speed at high speed," a rapid downward gesture corresponds to "decrease wind speed at high speed," and a rapid rotating gesture corresponds to "change wind direction at high speed." The control system establishes a mapping relationship between gestures and operation commands, achieving a direct correspondence between user gestures and smart fan operation commands. Simultaneously, a fuzzy logic controller is introduced to enhance the response capability to subtle differences in gestures, ensuring the accuracy and timeliness of operation commands. Finally, the system generates a smart fan operation control command corresponding to each specific user gesture action.
[0028] Step S4: Apply the smart fan operation control command corresponding to each user's specific gesture to the smart fan's built-in control system to execute the corresponding smart fan gesture action control working mode.
[0029] In this embodiment of the invention, the smart fan operation control command corresponding to each user's specific gesture action is transmitted to the smart fan's built-in control system to execute the corresponding gesture action control working mode. The control system is implemented based on a microcontroller (such as Arduino or Raspberry Pi). After receiving the command, it controls the motor speed through PWM (Pulse Width Modulation) signal to adjust the fan's working state. The system's internal feedback mechanism monitors the fan's operating state in real time, such as wind speed and direction, and compares it with the user's gesture command to ensure the accuracy and effectiveness of the fan operation.
[0030] Furthermore, as an embodiment of the present invention, referring to FIG2, which is a detailed flowchart of step S1 in FIG1, step S1 in this embodiment includes the following steps:
[0031] Step S11: Obtain a set of real-time gesture images of the user through the machine vision module built into the smart fan;
[0032] In this embodiment of the invention, a machine vision module built into the smart fan is used to capture real-time images of user gestures. This module is equipped with a high-resolution camera and an image processing unit, and can capture corresponding user gesture images at a rate of 30 frames per second. First, the user performs a gesture operation within a designated area in front of the fan. The machine vision module is then activated to acquire image data within that area in real time. To ensure image clarity and stability, automatic exposure and white balance technology are used to adapt to different lighting conditions. The acquired image data is preprocessed to form a set of gesture images, ensuring that the dynamic changes of the user's gestures can be accurately captured, ultimately resulting in a set of real-time user gesture images.
[0033] Step S12: Perform image size normalization processing on each frame of the user's real-time gesture action image set to obtain a user gesture action size normalized image set.
[0034] In this embodiment of the invention, the size of each frame of the user's real-time gesture image set obtained from previous real-time monitoring is normalized. By setting a target size, such as 640x480 pixels, each frame of the image is scaled using a bilinear interpolation algorithm to ensure that the image is minimized in the scaling process. The normalized image is then processed by grayscale to further improve the contrast and make the gesture edges more obvious. At this time, all pixel values of each frame of the gesture image vary between 0 and 255, which is convenient for subsequent edge detection processing, and finally, a user gesture size normalized image set is obtained.
[0035] Step S13: Obtain the set of non-zero pixels of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size, and perform gesture edge contour analysis on the corresponding gesture action image based on the set of non-zero pixels of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size, so as to generate the user gesture action edge contour region line corresponding to each frame of the gesture action image.
[0036] In this embodiment of the invention, the set of non-zero pixel points of the gesture edge of each frame of the gesture action image in the previously normalized image set of user gesture action size is obtained and labeled. These non-zero pixel point sets constitute the basic dataset of gesture edges. By combining the non-zero pixel point set of the gesture edge of each frame of the gesture action image with the Hough transform algorithm, the edge contour of the corresponding gesture action image is planned and connected to further analyze the edge contour, thereby generating the edge contour region line corresponding to each frame of the gesture action image. The contour region line clearly outlines the shape features of the gesture. Finally, the user gesture action edge contour region line corresponding to each frame of the gesture action image is generated.
[0037] Step S14: Based on the user gesture action edge contour region line corresponding to each frame of gesture action image, perform gesture background cropping processing on the corresponding gesture action image in the real-time user gesture action image set to obtain the user gesture action background cropped image set.
[0038] In this embodiment of the invention, the corresponding gesture action image in the real-time gesture action image set is cropped between the gesture area and the background area by combining the user gesture action edge contour area line corresponding to each frame of gesture action image obtained by previous planning. The outer rectangular area of the gesture is determined by the contour area line. Then, the original image is cropped using the rectangular area. This operation adopts an image cropping algorithm to retain only the gesture part and remove the background information. After cropping, the image is further sharpened to improve its contrast and edge details, ensuring that the cropped gesture image is clearly distinguishable. Finally, the user gesture action background cropped image set is obtained.
[0039] Step S15: Perform gesture recognition and temporal feature point extraction on the user gesture action background cropped image set to obtain the temporal change feature point set of each user gesture action, where user gesture actions include up, down and rotation gesture actions.
[0040] In this embodiment of the invention, specific gesture actions are identified and analyzed, and feature points under corresponding temporal changes are extracted from the previously cropped set of user gesture action background images. By using a deep learning model (such as a convolutional neural network) to extract features from the cropped images, the gesture actions can be identified. The model is trained on a large number of gesture samples and can accurately identify gesture actions such as up, down, and rotation. Based on the previously identified specific gesture actions, temporal change feature points are extracted to obtain the corresponding gesture parts temporal change feature points. The extracted feature points are arranged in chronological order to form a set of temporal feature points, representing the dynamic changes of gesture actions. Finally, a set of temporal change feature points of gesture parts corresponding to each user gesture action is obtained, where user gesture actions include up, down, and rotation gesture actions.
[0041] Furthermore, the step S13 of obtaining the set of non-zero pixel points of the gesture edge of each frame of the gesture action image within the normalized image set of the user gesture action size includes the following steps:
[0042] For each frame of the gesture action image in the normalized image set of user gesture action size, non-zero edge pixel points are detected and extracted to obtain the initial coordinate point set of non-zero pixel points of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size.
[0043] In this embodiment of the invention, each frame of the gesture action image in the normalized image set of user gesture action size is preprocessed using an image processing tool (such as OpenCV) to ensure that the contrast and brightness of the image are moderate, and the coordinates of all non-zero pixels in each frame of the gesture action image are obtained to form the initial coordinate point set of non-zero pixels at the gesture edge. This process is achieved by iterating through the pixels in each frame of the gesture action image and recording the positions of pixels with non-zero values, forming a two-dimensional coordinate list, and finally obtaining the initial coordinate point set of non-zero pixels at the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size.
[0044] Preferably, pixel density growth analysis is performed on each non-zero pixel point in the initial coordinate point set of non-zero pixels at the gesture edge of each frame of the gesture action image to obtain the pixel density around each non-zero pixel point in each frame of the gesture action image.
[0045] In this embodiment of the invention, statistical analysis of pixel density growth is performed on each non-zero pixel in the initial coordinate point set of non-zero pixels at the gesture edge of each frame of gesture action image previously detected and extracted. By defining a window size (e.g., a 3x3 or 5x5 neighborhood), and then calculating the total number of non-zero pixels in the neighborhood of each non-zero pixel, the pixel density around each non-zero pixel can be determined, forming a density map. This process can distinguish between high-density and low-density regions by setting a threshold, ensuring the accuracy of subsequent isolated point removal steps, and finally obtaining the pixel density around each non-zero pixel in each frame of gesture action image.
[0046] Preferably, based on the pixel density around each non-zero pixel in each frame of the gesture action image, an isolated point removal operation is performed on the corresponding non-zero pixel in the initial coordinate point set of the non-zero pixel of the gesture edge in each frame of the gesture action image to obtain the isolated removal point set of the non-zero pixel of the gesture edge in each frame of the gesture action image.
[0047] In this embodiment of the invention, by combining the pixel density around each non-zero pixel in each frame of the gesture action image obtained from previous analysis, an opening operation is performed on the corresponding non-zero pixel in the initial coordinate point set of non-zero pixels at the gesture edge to remove isolated points. This utilizes the opening operation in image morphology to process each frame of the gesture action image. The opening operation is composed of corresponding erosion operations to erode each non-zero pixel to remove isolated points in low-density areas, ensuring that only high-density non-zero pixels are retained. Finally, the set of isolated points for non-zero pixels at the gesture edge of each frame of the gesture action image is obtained.
[0048] Preferably, an edge hole filling and closing operation is performed on the set of isolated non-zero pixel points at the gesture edge of each frame of the gesture action image to obtain the set of non-zero pixel points at the gesture edge of each frame of the gesture action image within the user gesture action size normalized image set.
[0049] In this embodiment of the invention, a closing operation is performed to fill the non-zero pixel isolated removal point set of the gesture edge of each frame of gesture action image obtained from previous analysis. The purpose of this process is to enhance the overall shape of the gesture through morphological operations, making the edges more coherent. The closing operation is used to fill small holes in the image by first performing a dilation operation, and then performing an erosion operation to restore the features of the gesture edge. By selecting appropriate structuring elements, it is ensured that the holes of the edge can be filled without introducing new noise. In implementation, the closing operation function provided by the image processing library can be used to finally obtain the non-zero pixel point set of the gesture edge of each frame of gesture action image in the user gesture action size normalized image set.
[0050] Furthermore, step S15 includes the following steps:
[0051] Step S151: Perform quantitative motion vector analysis on each frame of the user gesture action image in the background cropping image set to obtain the gesture action motion direction vector and gesture action motion speed vector corresponding to each frame of the gesture action image.
[0052] In this embodiment of the invention, quantitative analysis of motion vectors is performed on each frame of the user gesture action background cropping image set obtained after background cropping. By collecting data from each frame of the image, the image sequence is processed using optical flow methods (such as the Lucas-Kanade method or the Horn-Schunck algorithm). This process involves calculating the pixel movement between adjacent frames, extracting motion information using sparse or dense optical flow methods, and tracking the pixels of each frame of the image to calculate the corresponding motion direction vector and motion velocity vector. The motion direction vector can be determined by calculating the angle of the motion vector, while the motion velocity is obtained by calculating the ratio of pixel change to time interval. Finally, the motion direction vector and motion velocity vector of each frame of the gesture action image are obtained.
[0053] Step S152: Based on the gesture motion direction vector and gesture motion speed vector corresponding to each frame of gesture motion image, perform gesture motion recognition and classification on the corresponding gesture motion images in the user gesture motion background cropping image set to obtain a subset of gesture motion background cropping images corresponding to each user gesture motion, wherein the user gesture motion includes up, down and rotation gesture motions.
[0054] In this embodiment of the invention, by combining the gesture motion direction vector and gesture motion speed vector corresponding to each frame of gesture motion image obtained from previous analysis, the corresponding gesture motion images in the user gesture motion background cropping image set are used to identify and classify user gesture motions, thereby defining a classifier, such as a support vector machine (SVM) or a convolutional neural network (CNN). This classifier uses the extracted motion vectors as input features. By comparing the motion features of each frame of image with a predefined gesture template, the classifier will determine the type of gesture, including gestures such as up, down, and rotation. During the training process, a labeled gesture dataset can be used for model training to improve the recognition accuracy. After the recognition is completed, the corresponding gesture motion images are summarized into a subset, and finally, the gesture motion background cropping image subset corresponding to each user gesture motion is obtained.
[0055] Step S153: Extract gesture feature points from each frame of the gesture action image within the subset of the gesture action background cropping image corresponding to each user's gesture action, to obtain the set of gesture feature points for each frame of the gesture action sub-image corresponding to each user's gesture action.
[0056] In this embodiment of the invention, the feature points of the hand gestures are extracted from each frame of the gesture action image within the previously identified and classified subset of user gesture action background images. The methods used include gesture recognition frameworks such as OpenPose or MediaPipe. These tools can efficiently extract the feature points of the hand. In specific operation, each frame of image is first input into the feature point extraction model, which returns the coordinates of the hand feature points, such as the thumb and index finger. Then, the extracted feature points are normalized to ensure that the feature points of different gestures are compared in the same coordinate system, thereby forming a set containing the feature point set corresponding to each frame of image. Finally, the feature point set of each user gesture action corresponding to each frame of gesture action sub-image is obtained.
[0057] Step S154: Perform gesture action interaction change analysis on each frame of gesture action image within the gesture action background cropping image subset corresponding to each user's gesture action to obtain the gesture action interaction change pattern between each frame of gesture action sub-image corresponding to each user's gesture action.
[0058] In this embodiment of the invention, the gesture interaction changes of each frame of the gesture action image within the previously identified and classified subset of user gesture actions are analyzed. In this process, algorithms such as Dynamic Time Warping (DTW) or Hidden Markov Model (HMM) are used to model the gesture actions and analyze the dynamic change patterns between different gestures. First, by performing temporal analysis on the extracted gesture feature point sequence, the movement patterns and change trends of each key point between different frames are identified. Then, by comparing the change trajectories of different gestures, the association and transformation relationships between gesture actions are identified, forming a description of the gesture action interaction change pattern. This helps to understand the dynamic behavior of users when performing specific gestures, and finally, the gesture action interaction change pattern between each frame of the user gesture action is obtained.
[0059] Step S155: Based on the gesture interaction change pattern between each user's gesture action corresponding to each frame's gesture action sub-image, perform temporal aggregation of the same gesture interaction process features of the gesture part feature set of each user's gesture action corresponding to each frame's gesture action sub-image to obtain the temporal change feature point set of each user's gesture action.
[0060] In this embodiment of the invention, the feature point set of each user's gesture action corresponding to each frame of the gesture action sub-image obtained by previous identification and analysis is integrated into the feature time sequence change point set within the same gesture interaction process by combining the gesture action interaction change pattern between each user's gesture action corresponding to each frame of the gesture action sub-image. In order to use aggregation techniques, such as weighted average or max pooling methods, the feature points under the same gesture interaction process are combined into a time sequence feature set. This process involves integrating the coordinates of each frame feature point according to the time sequence and using algorithms such as principal component analysis (PCA) for dimensionality reduction to reduce data redundancy while retaining key features so that they can effectively represent the user's gesture action. Finally, the feature point set of the gesture part time sequence change corresponding to each user's gesture action is obtained.
[0061] Furthermore, step S2 includes the following steps:
[0062] Step S21: Divide the set of temporal change feature points of each user's gesture action into a hierarchical representation of the gesture feature space, and obtain the subset of gesture feature points of each user's gesture action under different hierarchical representation positions.
[0063] In this embodiment of the invention, the gesture feature points corresponding to the temporal changes of each user gesture action obtained by previous aggregation are divided into different levels of representation in the gesture feature space. By using image processing technology and deep learning algorithms, the gesture feature points are divided into multiple levels according to the complexity and action characteristics of the user gesture action. The corresponding levels are divided into different levels such as fingers, palms, and forearms. Finally, a subset of gesture feature points of each user gesture action under different hierarchical representation positions is obtained.
[0064] Step S22: Perform a gesture importance assessment analysis on each gesture feature point in the subset of gesture feature points under different hierarchical representation positions for each user's gesture action, and obtain the degree of importance of each gesture feature point corresponding to each user's gesture action under different hierarchical representation positions; based on the degree of importance of each gesture feature point corresponding to each user's gesture action under different hierarchical representation positions, filter the key feature points of hierarchical representation to obtain the set of key feature points of gesture part hierarchical representation corresponding to each user's gesture action.
[0065] In this embodiment of the invention, the importance of each gesture feature point in the subset of gesture feature points at different hierarchical representation positions of each user gesture action is evaluated and calculated by sampling weighted analysis. This is to define evaluation indicators for each gesture feature point, including position stability, movement speed, and change frequency. Statistical analysis tools are used to calculate the contribution of each feature point throughout the gesture process, quantifying its impact on gesture recognition. For example, heatmaps are used to analyze the activity level of different feature points and assess their importance in gesture recognition. This yields the importance of each gesture action corresponding to each gesture feature point at different hierarchical representation positions of each user gesture action. Simultaneously, by combining the importance of each gesture feature point corresponding to each user gesture action at different hierarchical representation positions obtained from previous evaluations, key feature points of hierarchical representation are filtered out. By setting an importance threshold, feature points that significantly affect the recognition effect in gesture actions are selected. For example, setting the threshold to 0.5 filters out feature points with an importance evaluation below this value. During the filtering process, feature points that appear more frequently in multiple user gestures are preferentially retained. This can enhance the accuracy and robustness of gesture recognition, and finally obtain the set of key feature points of hierarchical representation of gesture parts corresponding to each user gesture action.
[0066] Step S23: Perform feature point vector transformation on the key feature point set of the gesture part hierarchical representation corresponding to each user's gesture action to obtain the key feature point vector of the gesture part hierarchical representation corresponding to each user's gesture action.
[0067] In this embodiment of the invention, after obtaining the set of key feature points representing the gesture parts of each user's gesture, the feature point vectors are transformed. This process uses a vectorization method to transform the coordinate values and statistics of each key feature point into spatial vectors, forming a unified feature vector. In specific implementation, linear transformation or normalization methods are used to ensure that all feature point vectors have the same scale, and the coordinate values of each feature point are combined according to the defined feature dimensions (such as x, y, z coordinates, speed, angle, etc.) to form a high-dimensional feature vector, finally obtaining the key feature point vectors representing the gesture parts of each user's gesture.
[0068] Step S24: Based on the key feature point vector of the gesture part hierarchical representation corresponding to each user's gesture action, the set of temporal change feature points of the gesture part corresponding to each user's gesture action is sorted and divided into similar gesture action pairs to obtain the set of temporal feature points of the gesture part corresponding to each similar gesture action.
[0069] In this embodiment of the invention, by combining the key feature point vectors of the gesture parts corresponding to each user gesture action obtained by previous conversion, a similarity calculation algorithm (such as cosine similarity or Euclidean distance) is used to compare each gesture vector and identify the gesture action sequence pairs with the same similarity. In this step, similar gesture actions are grouped by using clustering analysis methods, and the feature point set corresponding to the gesture actions in each group is filtered and divided to obtain the temporal feature point set of each similar gesture action corresponding to the gesture parts, and finally the temporal feature point set of each similar gesture action is obtained.
[0070] Step S25: Perform a detailed analysis of the differences between similar gestures based on the temporal feature point set of the gesture parts corresponding to each similar gesture, and obtain the detailed recognition results of the smart fan user gesture differences. The detailed recognition results of the smart fan user gesture differences include slightly upward, slightly downward, and slightly rotating waving gestures, as well as quickly upward, quickly downward, and quickly rotating waving gestures.
[0071] In this embodiment of the invention, the differences between similar gestures are refined by combining the temporal feature point sets of the gesture parts corresponding to the previously divided similar gestures. This process mainly analyzes the subtle differences between similar gestures. For example, for the two gestures of "slightly upward" and "quickly upward", the differences in the temporal changes of their key feature point vectors are compared. By conducting in-depth analysis of the waving rate and waving amplitude changes of the key feature point vectors corresponding to the gestures, the subtle differences in the gesture feature performance of the two types of gestures can be extracted, thereby identifying and dividing the gesture difference recognition results, covering slightly upward, slightly downward, and slightly rotating waving gestures as well as quickly upward, quickly downward, and quickly rotating waving gestures, and finally obtaining the refined recognition results of the smart fan user gesture differences.
[0072] Furthermore, step S21 includes the following steps:
[0073] The gesture feature structure space is transformed by performing a gesture feature structure space transformation on the set of gesture part temporal change feature points corresponding to each user's gesture action to obtain the gesture part temporal change feature space corresponding to each user's gesture action.
[0074] In this embodiment of the invention, by obtaining the set of temporal change feature points of each user's gesture action corresponding to the gesture position obtained from previous analysis, including the coordinate change of each feature point and its corresponding timestamp, and by using the transformation of the feature space coordinate system, the feature point set of the gesture position is spatially transformed by using feature extraction algorithms (such as principal component analysis, Fourier transform, etc.), thereby transforming the coordinates of the feature points in three-dimensional space into a unified feature space, and finally obtaining the temporal change feature space of each user's gesture action corresponding to the gesture position.
[0075] Preferably, the time distribution and motion amplitude of each gesture feature point in the temporal change feature space of each user's gesture action are analyzed to obtain the time distribution and motion change amplitude data of each gesture feature point in the temporal change feature space of each user's gesture action.
[0076] In this embodiment of the invention, a detailed analysis is performed on each feature point of the gesture part in the temporal change feature space corresponding to the converted user gesture action. By using time series analysis methods, the time distribution characteristics of each feature point are modeled, and the state of the feature point at each timestamp during the gesture execution process is calculated. At the same time, by calculating the motion amplitude of the feature point, its displacement change in different time periods is analyzed. A dynamic time warping algorithm can be used to process the time series data, thereby obtaining the time distribution and specific motion amplitude of each gesture part feature point in the temporal change process. Finally, the time distribution and motion change amplitude data corresponding to each gesture part feature point in the temporal change feature space of each user gesture action are obtained.
[0077] Preferably, based on the time distribution and motion change amplitude data of each gesture feature point in the temporal change feature space of each user's gesture action, a hierarchical representation mapping analysis is performed on each gesture feature point in the temporal change feature space of each user's gesture action to obtain the hierarchical structure representation mapping positional relationship of each gesture feature point in the temporal change feature space of each user's gesture action.
[0078] In this embodiment of the invention, by combining the time distribution and motion change amplitude data corresponding to each gesture feature point in the temporal change feature space of each user's gesture action obtained from previous statistical analysis, a hierarchical spatial representation mapping analysis is performed on each gesture feature point in the corresponding temporal change feature space of the gesture part. In order to use a hierarchical clustering algorithm, the gesture feature points are grouped to form a hierarchical structure representation, and the hierarchical structure positional relationship of each feature point is determined. Finally, the hierarchical structure representation mapping positional relationship of each gesture feature point in the temporal change feature space of each user's gesture action is obtained.
[0079] Preferably, the corresponding gesture feature points are hierarchically represented and divided according to the hierarchical structure representation mapping position relationship of each gesture feature point in the temporal change feature space of each user's gesture action, so as to obtain the gesture feature point position subset of each user's gesture action under different hierarchical representation positions.
[0080] In this embodiment of the invention, the corresponding gesture feature points in each user's gesture action are hierarchically represented by combining the hierarchical structure representation of gesture feature points obtained from previous statistical analysis. In the specific implementation process, the gesture feature points are first divided into different levels, such as the wrist layer, finger layer, and elbow layer, etc., and the hierarchical representation at different positions is performed. The clustering analysis results are used for classification, and the feature points are filtered by setting a threshold. Feature points with similar features and movement patterns are assigned to the same level and form a subset of gesture feature points. Finally, the subset of gesture feature points in different hierarchical representation positions of each user's gesture action is obtained.
[0081] Furthermore, step S24 includes the following steps:
[0082] Step S241: Use a multi-layer convolutional neural network to perform different level representation feature division on the key feature point vectors of the gesture parts corresponding to each user's gesture action, and obtain the network layer of key feature point vectors of different gesture parts corresponding to each user's gesture action.
[0083] In this embodiment of the invention, a multi-layer convolutional neural network (CNN) is used to divide the key feature point vectors of the gesture parts corresponding to each user's gesture action into different hierarchical representations. The key feature point vectors of the gesture parts hierarchical representation obtained in the previous analysis are input into the multi-layer convolutional neural network. Each layer of the network contains multiple convolutional kernels. These convolutional kernels extract the feature point vectors of different levels of the gesture parts hierarchical representation through convolution operations. These feature point vectors in each level of the network not only retain the spatial information of key feature points at different gesture spatial parts, but also contain information on the dynamic changes over time. Finally, different key feature point vector network layers corresponding to each user's gesture action are obtained.
[0084] Step S242: Calculate the cosine similarity between the key feature point vectors of different gesture parts in the network layer corresponding to each user's gesture action, so as to obtain the cosine similarity of the feature point vectors of each user's gesture action between different gesture parts in the network layer.
[0085] In this embodiment of the invention, the cosine similarity calculation method is used to quantify the cosine similarity between the key feature point vectors of each user gesture action in different gesture part key feature point vector network layers obtained in the previous division. First, the key feature point vectors of each user gesture action in the same level representation network layer need to be extracted to form multiple vector subsets. Then, the cosine similarity between the feature point vectors of any two user gesture actions in the same network layer is quantified. This similarity is obtained by calculating the dot product of the two vectors and dividing it by their modulus product. Finally, the cosine similarity of the feature point vectors of each user gesture action between different gesture part network layers is obtained.
[0086] Step S243: Calculate the cosine similarity arithmetic mean of the feature point vectors of each user's gesture action between different gesture parts network layers to obtain the average similarity of the feature point vectors of each user's gesture action.
[0087] In this embodiment of the invention, the cosine similarity of the feature point vectors of each user gesture action between different gesture parts network layers, which was previously quantized and calculated in different network layers, is arithmetically averaged to obtain the average similarity of the feature point vectors of each user gesture action. This step ensures that the similarity of different user gesture actions can be represented by a unified value, and finally the average similarity of the feature point vectors of each user gesture action is obtained.
[0088] Step S244: Based on the average similarity of the feature point vectors of the gesture parts between each user's gesture actions, perform similar gesture action pair filtering on the corresponding user gesture actions to obtain similar gesture action sequence pairs.
[0089] In this embodiment of the invention, the corresponding user gesture actions are filtered for similar gesture action pairs by combining the average similarity of the gesture feature point vectors between each user gesture action obtained by the previous average calculation. By setting a threshold, it is determined which gesture actions can be considered "similar". Then, based on the previously calculated average similarity, the gesture action pairs that meet the threshold condition are extracted to form similar gesture action sequence pairs of users. This process adopts an automated method. By writing a specific algorithm to compare and filter, it is ensured that the obtained gesture sequence pairs have a high degree of consistency at the feature level, and finally, similar gesture action sequence pairs of users are obtained.
[0090] Step S245: Based on the user's similar gesture action sequence, divide the gesture part temporal change feature point set corresponding to each user's gesture action into similar gesture action feature point set, and obtain the gesture part temporal feature point set corresponding to each similar gesture action.
[0091] In this embodiment of the invention, by combining the previously selected user similar gesture action sequence with the corresponding similar gesture action pairs, the corresponding gesture action temporal change feature point set of the gesture parts is matched and divided to extract the feature point set corresponding to each similar gesture action. The feature point set is then grouped according to the change pattern of the time series. This process ensures that each similar gesture action is not only similar in feature space, but also consistent in temporal evolution. Finally, the temporal feature point set of each similar gesture action corresponding to the gesture parts is obtained.
[0092] Furthermore, step S25 includes the following steps:
[0093] Step S251: Map the hand gesture waving trajectory of the temporal feature point set of each similar hand gesture to generate the hand gesture waving trajectory of each similar hand gesture.
[0094] In this embodiment of the invention, the gesture waving trajectory is mapped by the temporal feature point set (including the coordinates of key feature points such as wrist, fingers, elbow and shoulder) corresponding to various similar gestures obtained by previous analysis. The waving trajectory is smoothed by interpolation algorithm to eliminate the jitter caused by the movement, thereby obtaining the corresponding real-time gesture waving trajectory. These trajectories will be standardized to the same time length, and finally the gesture waving trajectory corresponding to various similar gestures will be generated.
[0095] Step S252: Perform hand gesture waving time frame segmentation on the waving trajectory of the hand gesture part corresponding to each similar hand gesture to obtain the hand gesture waving time frame trajectory segmentation corresponding to each similar hand gesture.
[0096] In this embodiment of the invention, the waving trajectory of the hand gesture corresponding to each similar gesture generated by previous mapping is divided into time frame segments. By setting an appropriate time frame window (e.g., 100 milliseconds), the instantaneous state of each gesture is captured. Subsequently, based on the characteristics of the waving trajectory, the waving trajectory is segmented using the Dynamic Time Warping (DTW) algorithm, dividing the continuous waving action into multiple time frame segments. Each time frame segment contains corresponding feature point data and is saved in chronological order. This process aims to clearly reflect the changing feature trajectory segmentation of the gesture in different time periods, and finally obtains the waving time frame trajectory segments of each similar gesture.
[0097] Step S253: Perform inter-frame waving rate and amplitude change gradient analysis on the waving time frame trajectory segments corresponding to each similar gesture to obtain the waving rate change gradient and the amplitude change gradient of the gesture between each similar gesture time frame.
[0098] In this embodiment of the invention, mathematical statistical tools are used to perform gradient statistical calculations on the rate and amplitude changes of the hand gesture waving time frame trajectories corresponding to the previously segmented similar hand gestures. By calculating the change in the position of the hand gesture feature points in adjacent time frames, the waving rate of each time frame is obtained. The rate calculation formula is: rate = position change / time interval. Then, the amplitude change of each time frame is calculated. The amplitude is defined as the distance between the maximum and minimum feature points of the hand gesture trajectory. Gradient analysis is performed on these data, and the difference calculation method is used to evaluate the changes in rate and amplitude. The gradient of the hand gesture waving rate change and the gradient of the hand gesture waving amplitude change between each time frame are identified. Finally, the gradient of the hand gesture waving rate change and the gradient of the hand gesture waving amplitude change between the time frames of each similar hand gesture are obtained.
[0099] Step S254: Based on the gradient of the change in the waving rate and the gradient of the change in the waving amplitude of the waving action between time frames of various similar gesture actions, the waving action waving subtle difference calculation formula is used to segment the waving action waving time frame trajectory corresponding to each similar gesture action to calculate the waving subtle difference distinction between each similar gesture action.
[0100] In this embodiment of the invention, a suitable formula for calculating subtle differences in gesture waving is constructed by combining measurement parameters of similar gestures, time variable parameters, waving rate of each similar gesture, gradient of waving rate change, waving rate weighting coefficient, waving amplitude of each similar gesture, gradient of waving amplitude change, waving amplitude weighting coefficient, waving time difference attenuation coefficient, and related parameters. This formula is used to segment the waving time frame trajectory of each similar gesture to distinguish subtle differences in waving. The weighted difference method is then used to compare and calculate the trajectories of each time frame of the similar gestures, quantifying the subtle differences between different gestures, and finally obtaining the distinguishability of subtle differences in waving between each similar gesture.
[0101] Step S255: Based on the discriminative power of subtle differences in gesture waving between various similar gesture actions, perform a detailed analysis of similar gesture differences for the corresponding similar gesture actions to obtain the detailed recognition results of smart fan user gesture differences. The detailed recognition results of smart fan user gesture differences include slightly upward, slightly downward, and slightly rotating waving gesture actions as well as quickly upward, quickly downward, and quickly rotating waving gesture actions.
[0102] In this embodiment of the invention, a detailed analysis of the corresponding similar gestures is performed by combining the distinguishability of subtle differences in gesture waving between various similar gestures obtained through previous quantitative calculations. At this time, the distinguishability of subtle differences in gesture waving obtained through the above quantitative calculations is compared and judged according to a preset waving difference threshold. If the distinguishability of subtle differences in gesture waving is less than the preset waving difference threshold, the corresponding similar gesture is identified as a slight state, such as a slight upward, slight downward, or slight rotational waving action. If the distinguishability of subtle differences in gesture waving is greater than or equal to the preset waving difference threshold, the corresponding similar gesture is identified as a fast state, such as a fast upward, fast downward, or fast rotational waving action. The recognition result of gesture differences is then output to ensure that the user's gestures can be accurately recognized and converted into corresponding fan control commands. Finally, the refined recognition result of the smart fan user gesture differences is obtained, which includes slight upward, slight downward, and slight rotational waving gestures as well as fast upward, fast downward, and fast rotational waving gestures.
[0103] Furthermore, the formula for calculating subtle differences in hand gesture movements in step S254 is as follows:
[0104] In the formula, D ij V represents the distinguishability of subtle differences in hand gestures between the i-th and j-th similar hand gestures, where i and j are both measurement parameters for similar hand gestures, t1 is the start time of the time interval, t2 is the end time of the time interval, t is the integration time variable parameter, and V i (t) represents the waving speed of the i-th similar gesture at time t, V j (t) represents the waving rate of the j-th similar gesture at time t. Let A be the gradient of the waving rate change between the i-th and j-th similar hand gestures, and α be the waving rate weighting coefficient. i (t) represents the amplitude of the waving gesture of the i-th similar gesture at time t, A j (t) represents the amplitude of the hand gesture at time t for the j-th similar hand gesture. Let β be the gradient of the amplitude change of the gesture between the i-th similar gesture and the j-th similar gesture, λ be the amplitude weighting coefficient, λ be the gesture wave time difference attenuation coefficient, and η be the correction coefficient for the subtle difference in gesture wave distinction.
[0105] This invention, through the use of a specific mathematical model and verification, derives a formula for calculating subtle differences in gesture waving. This formula is used to segment the waving time frame trajectory of similar gestures and differentiate between subtle differences. It allows for precise quantification of subtle differences between different gestures. By calculating changes in waving rate and amplitude, it can more accurately identify user intentions, thereby improving the responsiveness and interactive experience of smart devices. The formula incorporates information from multiple dimensions, such as waving rate and amplitude, and uses time decay to consider the time factor, enhancing its adaptability to gestures and ensuring that outdated information is effectively ignored, focusing on the current relevant action. Secondly, the weighting coefficients in the formula provide flexibility in the calculation process, allowing the importance of rate and amplitude to be adjusted according to actual needs in specific application scenarios. Furthermore, the introduction of correction coefficients can consider other external factors that may affect gesture recognition, such as environmental noise or individual user differences, thereby improving the reliability of the calculation results. This mechanism enhances the robustness of the calculation formula, ensuring stable gesture recognition even in complex environments. Therefore, this formula fully considers the subtle differences in gesture waving between the i-th and j-th similar gestures, specifically the discriminative power D. ij Similar to the gesture, the measurement parameters i and j, the start time point t1 of the time interval, the end time point t2 of the time interval, the integration time variable parameter t, and the waving rate V of the i-th similar gesture at time t. i (t), the waving rate V of the j-th similar gesture at time t. j (t), the gradient of the waving rate change between the i-th and j-th similar gestures. The waving rate weighting coefficient α, and the waving amplitude A of the i-th similar gesture at time t. i (t), the amplitude A of the j-th similar gesture at time t. j (t), the gradient of the amplitude change of the gesture movement between the i-th similar gesture and the j-th similar gesture. The waving amplitude weighting coefficient β, the gesture waving time difference attenuation coefficient λ, the correction coefficient η for the subtle difference discrimination of gesture waving, and the subtle difference discrimination D between the i-th and j-th similar gestures. ij The interrelationships between the above parameters constitute a functional relationship. This formula can distinguish the subtle differences in the waving time frame trajectory of various similar gestures. At the same time, by introducing the correction coefficient η of the subtle difference in waving, it can be adjusted according to the error that occurs in the calculation process, thereby improving the accuracy and applicability of the formula for calculating the subtle differences in waving.
[0106] Furthermore, the present invention also provides a machine vision-based intelligent fan gesture recognition system for executing the machine vision-based intelligent fan gesture recognition method described above. The machine vision-based intelligent fan gesture recognition system includes:
[0107] The gesture action temporal feature point extraction module is used to acquire a set of real-time user gesture action images through the machine vision module built into the smart fan, and to perform gesture background cropping processing on the set of real-time user gesture action images to obtain a set of user gesture action background cropped images; and to perform gesture action recognition and temporal feature point extraction on the set of user gesture action background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture action, wherein the user gesture actions include upward, downward and rotation gesture actions;
[0108] The similar gesture action difference refinement and recognition module is used to transform the gesture-level key feature point vector of the gesture part temporal change feature point set corresponding to each user's gesture action to obtain the gesture part hierarchical representation key feature point vector of each user's gesture action; based on the gesture part hierarchical representation key feature point vector of each user's gesture action, similar gesture action pairs are filtered and divided into similar gesture action pairs to obtain the gesture part temporal feature point set corresponding to each similar gesture action; based on the gesture part temporal feature point set corresponding to each similar gesture action, similar gesture difference refinement analysis is performed to obtain the smart fan user gesture difference refinement and recognition results, which include slightly upward, slightly downward, slightly rotating waving gesture actions and quickly upward, quickly downward, quickly rotating waving gesture actions.
[0109] The user gesture refinement action control response module is used to perform gesture action operation control response for each specific action of the user gesture within the refined recognition result of the user gesture difference of the smart fan, so as to generate the smart fan operation control command corresponding to each specific action of the user gesture.
[0110] The user gesture control command response and execution module is used to apply the smart fan operation control command corresponding to each user gesture to the built-in control system of the smart fan, so as to execute the corresponding smart fan gesture action control working mode.
[0111] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. The present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A machine vision-based intelligent fan gesture recognition method, characterized in that, Includes the following steps: Step S1: Obtain a set of real-time user gesture images through the machine vision module built into the smart fan, and perform gesture background cropping on the set of real-time user gesture images to obtain a set of user gesture background cropped images; perform gesture recognition and temporal feature point extraction on the set of user gesture background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture, wherein the user gestures include up, down and rotation gestures. Step S2: Perform gesture-level key feature point vector transformation on the gesture part temporal change feature point set corresponding to each user's gesture action to obtain the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action; based on the gesture part hierarchical representation key feature point vector corresponding to each user's gesture action, perform similar gesture action pair filtering and division on the gesture part temporal change feature point set corresponding to each similar gesture action to obtain the gesture part temporal feature point set corresponding to each similar gesture action; perform similar gesture difference refinement analysis on the gesture part temporal feature point set corresponding to each similar gesture action to obtain the smart fan user gesture difference refinement recognition result, which includes slightly upward, slightly downward, slightly rotating waving gesture actions as well as fast upward, fast downward, fast rotating waving gesture actions; Step S3: Perform gesture operation control response for each user gesture in the refined recognition result of the smart fan user gesture differences, so as to generate the smart fan operation control command corresponding to each user gesture. Step S4: Apply the smart fan operation control command corresponding to each user's specific gesture to the smart fan's built-in control system to execute the corresponding smart fan gesture action control working mode.
2. The intelligent fan gesture recognition method based on machine vision according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Obtain a set of real-time gesture images of the user through the machine vision module built into the smart fan; Step S12: Perform image size normalization processing on each frame of the user's real-time gesture action image set to obtain a user gesture action size normalized image set. Step S13: Obtain the set of non-zero pixels of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size, and perform gesture edge contour analysis on the corresponding gesture action image based on the set of non-zero pixels of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size, so as to generate the user gesture action edge contour region line corresponding to each frame of the gesture action image. Step S14: Based on the user gesture action edge contour region line corresponding to each frame of gesture action image, perform gesture background cropping processing on the corresponding gesture action image in the real-time user gesture action image set to obtain the user gesture action background cropped image set. Step S15: Perform gesture recognition and temporal feature point extraction on the user gesture action background cropped image set to obtain the temporal change feature point set of each user gesture action, where user gesture actions include up, down, and... Rotating hand gesture.
3. The intelligent fan gesture recognition method based on machine vision according to claim 2, characterized in that, Step S13, which involves obtaining the set of non-zero pixels at the gesture edge of each frame of the gesture action image within the normalized image set of the user's gesture action size, includes the following steps: For each frame of the gesture action image in the normalized image set of user gesture action size, non-zero edge pixel points are detected and extracted to obtain the initial coordinate point set of non-zero pixel points of the gesture edge of each frame of the gesture action image in the normalized image set of user gesture action size. For each non-zero pixel in the initial coordinate point set of non-zero pixels at the gesture edge of each frame of the gesture action image, pixel density growth analysis is performed to obtain the pixel density around each non-zero pixel in each frame of the gesture action image. Based on the pixel density around each non-zero pixel in each frame of the gesture action image, the isolated point removal operation is performed on the corresponding non-zero pixel in the initial coordinate point set of the non-zero pixel of the gesture edge of each frame of the gesture action image to obtain the isolated point set of the non-zero pixel of the gesture edge of each frame of the gesture action image. For each frame of the gesture action image, the set of isolated non-zero pixels at the gesture edge is removed and the edge hole filling operation is performed to obtain the set of non-zero pixels at the gesture edge of each frame of the gesture action image within the user gesture action size normalized image set.
4. The intelligent fan gesture recognition method based on machine vision according to claim 2, characterized in that, Step S15 includes the following steps: Step S151: Perform quantitative motion vector analysis on each frame of the user gesture action image in the background cropping image set to obtain the gesture action motion direction vector and gesture action motion speed vector corresponding to each frame of the gesture action image. Step S152: Based on the gesture motion direction vector and gesture motion speed vector corresponding to each frame of gesture motion image, perform gesture motion recognition and classification on the corresponding gesture motion images in the user gesture motion background cropping image set to obtain a subset of gesture motion background cropping images corresponding to each user gesture motion, wherein the user gesture motion includes up, down and rotation gesture motions. Step S153: Extract gesture feature points from each frame of the gesture action image within the subset of the gesture action background cropping image corresponding to each user's gesture action, to obtain the set of gesture feature points for each frame of the gesture action sub-image corresponding to each user's gesture action. Step S154: Perform gesture action interaction change analysis on each frame of gesture action image within the gesture action background cropping image subset corresponding to each user's gesture action to obtain the gesture action interaction change pattern between each frame of gesture action sub-image corresponding to each user's gesture action. Step S155: Based on the gesture interaction change pattern between each user's gesture action corresponding to each frame's gesture action sub-image, perform temporal aggregation of the same gesture interaction process features of the gesture part feature set of each user's gesture action corresponding to each frame's gesture action sub-image to obtain the temporal change feature point set of each user's gesture action.
5. The intelligent fan gesture recognition method based on machine vision according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Divide the set of temporal change feature points of each user's gesture action into a hierarchical representation of the gesture feature space, and obtain the subset of gesture feature points of each user's gesture action under different hierarchical representation positions. Step S22: Perform a gesture importance assessment analysis on each gesture feature point in the subset of gesture feature points under different hierarchical representation positions for each user's gesture action, and obtain the degree of importance of each gesture feature point corresponding to each user's gesture action under different hierarchical representation positions; based on the degree of importance of each gesture feature point corresponding to each user's gesture action under different hierarchical representation positions, filter the key feature points of hierarchical representation to obtain the set of key feature points of gesture part hierarchical representation corresponding to each user's gesture action. Step S23: Perform feature point vector transformation on the key feature point set of the gesture part hierarchical representation corresponding to each user's gesture action to obtain the key feature point vector of the gesture part hierarchical representation corresponding to each user's gesture action. Step S24: Based on the key feature point vector of the gesture part hierarchical representation corresponding to each user's gesture action, the set of temporal change feature points of the gesture part corresponding to each user's gesture action is sorted and divided into similar gesture action pairs to obtain the set of temporal feature points of the gesture part corresponding to each similar gesture action. Step S25: Perform a detailed analysis of the differences between similar gestures based on the temporal feature point set of the gesture parts corresponding to each similar gesture, and obtain the detailed recognition results of the smart fan user gesture differences. The detailed recognition results of the smart fan user gesture differences include slightly upward, slightly downward, and slightly rotating waving gestures, as well as quickly upward, quickly downward, and quickly rotating waving gestures.
6. The intelligent fan gesture recognition method based on machine vision according to claim 5, characterized in that, Step S21 includes the following steps: The gesture feature structure space is transformed by performing a gesture feature structure space transformation on the set of gesture part temporal change feature points corresponding to each user's gesture action to obtain the gesture part temporal change feature space corresponding to each user's gesture action. The temporal distribution and motion amplitude of each gesture feature point in the temporal change feature space of each user's gesture action are analyzed to obtain the temporal distribution and motion change amplitude data of each gesture feature point in the temporal change feature space of each user's gesture action. Based on the time distribution and motion change amplitude data of each gesture feature point in the temporal change feature space of each user's gesture action, a hierarchical representation mapping analysis is performed on each gesture feature point in the temporal change feature space of each user's gesture action to obtain the hierarchical structure representation mapping position relationship of each gesture feature point in the temporal change feature space of each user's gesture action. Based on the hierarchical structure representation of the gesture feature points corresponding to each gesture part in the temporal variation feature space of each user's gesture actions, the corresponding gesture feature points are hierarchically represented and divided to obtain the hierarchical representation of each user's gesture actions in the time-series variation feature space. Subsets of gesture feature points under different hierarchical representations.
7. The intelligent fan gesture recognition method based on machine vision according to claim 5, characterized in that, Step S24 includes the following steps: Step S241: Use a multi-layer convolutional neural network to perform different level representation feature division on the key feature point vectors of the gesture parts corresponding to each user's gesture action, and obtain the network layer of key feature point vectors of different gesture parts corresponding to each user's gesture action. Step S242: Calculate the cosine similarity between the key feature point vectors of different gesture parts in the network layer corresponding to each user's gesture action, so as to obtain the cosine similarity of the feature point vectors of each user's gesture action between different gesture parts in the network layer. Step S243: Calculate the cosine similarity arithmetic mean of the feature point vectors of each user's gesture action between different gesture parts network layers to obtain the average similarity of the feature point vectors of each user's gesture action. Step S244: Based on the average similarity of the feature point vectors of the gesture parts between each user's gesture actions, perform similar gesture action pair filtering on the corresponding user gesture actions to obtain similar gesture action sequence pairs. Step S245: Based on the user's similar gesture action sequence, divide the gesture part temporal change feature point set corresponding to each user's gesture action into similar gesture action feature point set, and obtain the gesture part temporal feature point set corresponding to each similar gesture action.
8. The intelligent fan gesture recognition method based on machine vision according to claim 5, characterized in that, Step S25 includes the following steps: Step S251: Map the hand gesture waving trajectory of the temporal feature point set of each similar hand gesture to generate the hand gesture waving trajectory of each similar hand gesture. Step S252: Perform hand gesture waving time frame segmentation on the waving trajectory of the hand gesture part corresponding to each similar hand gesture to obtain the hand gesture waving time frame trajectory segmentation corresponding to each similar hand gesture. Step S253: Perform inter-frame waving rate and amplitude change gradient analysis on the waving time frame trajectory segments corresponding to each similar gesture to obtain the waving rate change gradient and the amplitude change gradient of the gesture between each similar gesture time frame. Step S254: Based on the gradient of the change in the waving rate and the gradient of the change in the waving amplitude of the waving action between time frames of various similar gesture actions, the waving action waving subtle difference calculation formula is used to segment the waving action waving time frame trajectory corresponding to each similar gesture action to calculate the waving subtle difference distinction between each similar gesture action. Step S255: Based on the discriminative power of subtle differences in gesture movements among similar gestures, perform a detailed analysis of the differences in similar gestures to obtain the detailed recognition results of the smart fan user's gesture differences. The refined recognition results of user gesture differences include slight upward, slight downward, and slight rotating waving gestures, as well as rapid upward, rapid downward, and rapid rotating waving gestures.
9. The intelligent fan gesture recognition method based on machine vision according to claim 8, characterized in that, The formula for calculating subtle differences in hand gestures in step S254 is as follows: In the formula, D ij V represents the distinguishability of subtle differences in hand gestures between the i-th and j-th similar hand gestures, where i and j are both measurement parameters for similar hand gestures, t1 is the start time of the time interval, t2 is the end time of the time interval, t is the integration time variable parameter, and V i (t) represents the waving speed of the i-th similar gesture at time t, V j (t) represents the waving rate of the j-th similar gesture at time t. Let A be the gradient of the waving rate change between the i-th and j-th similar hand gestures, and α be the waving rate weighting coefficient. i (t) represents the amplitude of the waving gesture of the i-th similar gesture at time t, A j (t) represents the amplitude of the hand gesture at time t for the j-th similar hand gesture. Let β be the gradient of the amplitude change of the gesture between the i-th similar gesture and the j-th similar gesture, λ be the amplitude weighting coefficient, λ be the gesture wave time difference attenuation coefficient, and η be the correction coefficient for the subtle difference in gesture wave distinction.
10. A machine vision-based intelligent fan gesture recognition system, characterized in that, For performing the machine vision-based intelligent fan gesture recognition method as described in claim 1, the machine vision-based intelligent fan gesture recognition system includes: The gesture action temporal feature point extraction module is used to acquire a set of real-time user gesture action images through the machine vision module built into the smart fan, and to perform gesture background cropping processing on the set of real-time user gesture action images to obtain a set of user gesture action background cropped images; and to perform gesture action recognition and temporal feature point extraction on the set of user gesture action background cropped images to obtain a set of temporal change feature points of gesture parts corresponding to each user gesture action, wherein the user gesture actions include upward, downward and rotation gesture actions; The similar gesture action difference refinement and recognition module is used to transform the gesture-level key feature point vector of the temporal change feature point set corresponding to each user's gesture action to obtain the gesture-level representation key feature point vector of each user's gesture action; based on the gesture-level representation key feature point vector of each user's gesture action, the module filters and divides the gesture-level feature point set corresponding to each user's gesture action into similar gesture action pairs to obtain the gesture-level temporal feature point set corresponding to each similar gesture action; based on the gesture-level temporal feature point set corresponding to each similar gesture action, the module performs similar gesture difference refinement analysis to obtain the smart fan user gesture difference refinement and recognition results, which include slightly upward, slightly downward, slightly rotating waving gestures, and rapid... Upward, rapid downward, and rapid rotating waving gestures; The user gesture refinement action control response module is used to perform gesture action operation control response for each specific action of the user gesture within the refined recognition result of the user gesture difference of the smart fan, so as to generate the smart fan operation control command corresponding to each specific action of the user gesture. The user gesture control command response and execution module is used to apply the smart fan operation control command corresponding to each user gesture to the built-in control system of the smart fan, so as to execute the corresponding smart fan gesture action control working mode.