Gesture somatosensory interaction method based on fused vision, inertial touch and multi-point touch information
By combining edge detection with three-frame difference algorithm and Gaussian mixture background difference algorithm, and using generative adversarial network model for gesture recognition, the problem of image contour information loss in gesture recognition is solved, and higher recognition rate and lower false detection rate are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG SPORT UNIV
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-08
AI Technical Summary
In the process of gesture recognition, the loss of image contour information leads to low recognition rate and high false detection rate, which is difficult to solve effectively with existing technologies.
By combining edge detection and three-frame difference algorithms with Gaussian mixture background subtraction algorithm, gesture recognition and classification are performed through generative adversarial network model, thereby improving the continuity of image contours and the integrity of internal information.
It improved the gesture recognition rate, reduced the false detection rate, and obtained images with continuous contours and complete internal information.
Smart Images

Figure CN121996065A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of gesture and motion-sensing interaction technology, specifically relating to a gesture and motion-sensing interaction method based on the fusion of visual, inertial tactile, and multi-touch information. Background Technology
[0002] The convergence of new-generation information technology, artificial intelligence, and advanced manufacturing technology has made human-computer interaction an important component in the development of intelligent equipment and products. In recent years, natural human-computer interaction technologies such as voice, gestures, postures, and facial expressions have endowed computing systems with strong perception capabilities, multi-channel capabilities, and naturalness, replacing the mouse and keyboard and becoming the interactive medium for human-machine communication, decision-making, and execution. In human-computer interaction scenarios, integrating gesture recognition technology to establish a dialogue mechanism between humans and machines will bring a more flexible and efficient experience. However, during the gesture recognition process, the obtained image contour information is often lost, and hollow displays appear, resulting in a high recognition rate and a high degree of false detection. Summary of the Invention
[0003] To address the technical problems mentioned above, this invention provides a gesture-based haptic interaction method that integrates visual, inertial, and multi-touch information. This method enables the acquisition of images with continuous contours and relatively complete internal information, thereby improving the gesture recognition rate and reducing the false detection rate.
[0004] In a first aspect, the present invention provides a gesture-based haptic interaction method based on the fusion of visual, inertial haptic, and multi-touch information, comprising: Step 1: Acquire video image sequence data of the hand; Step 2: Process the acquired video image sequence data to obtain the target image; Step 3: Perform gesture recognition and classification on the obtained target image using a generative adversarial network model, and send the recognition and classification results to the human-computer interaction module. The human-computer interaction interface displays the received video image sequence data and displays text and synthesizes speech based on the received recognition results.
[0005] Furthermore, step two includes: Acquire the (K-1), K, and K+1th frames of the video image sequence data, and perform grayscale conversion and median filtering on the acquired (K-1), K, and K+1th frames to obtain the video images respectively. , , ; The foreground target image is obtained by performing a background difference mixture Gaussian model on the (K+1)th frame of the video image. ; For video images and video images Perform differential processing to obtain the differential image. At the same time, for video images and video images Perform differential processing to obtain the differential image. ; For difference images Image obtained by performing Canny edge detection Simultaneously, for the difference image and Perform a logical AND operation to obtain the image. ; Image and images Perform a logical OR operation to obtain the image. ; Image and images Perform a logical AND operation to obtain a target image res with a clearer outline and internal information.
[0006] Furthermore, the generative adversarial network model includes an encoder, a decoder, a discriminator, and a classifier; The encoder is used to process the input target image. Compressed mapping to the latent space extracts the original features of the target image. ; The decoder takes the original features z of the target image and the one-hot encoding of the label as input, decodes them, and outputs the decoded image. ; The discriminator is used to classify the target image. and decoded image The input is processed by the discriminator, and the result is output by the Sigmoid layer as an arbitrary probability value between 0 and 1. The classifier distinguishes the target image. The category is represented by the probability of each category. The corresponding category can be determined by taking the index of the maximum value of the Argmax function's output.
[0007] Furthermore, the encoder includes four 3×3 convolutional layers with a stride of 2, two InceptionV2 structures, and one fully connected layer. The convolutional layers are equipped with CBN layers and Mish activation functions to normalize the output and enhance the network's expressive power.
[0008] Furthermore, the decoder includes a fully connected layer, four transposed convolutional networks, and two InceptionV2-trans networks. Among the four transposed convolutional networks, three layers use the Mish activation function, one layer uses the Tanh activation function, and CBN processing is performed before using the activation functions to prevent overfitting.
[0009] Furthermore, the discriminator consists of four convolutional layers and two fully connected layers, with the convolutional layers using the Mish activation function and the fully connected layers using the Sigmoid function.
[0010] Secondly, the present invention provides a computer-readable storage medium including a stored program that, when the program is running, controls the electrical equipment where the computer-readable storage medium is located to execute the gesture-based somatosensory interaction method described above, which integrates visual, inertial tactile, and multi-touch information.
[0011] The beneficial effects of this invention are as follows: This invention combines edge detection with a three-frame difference algorithm, resulting in images with excellent contour information and mitigating some "hole" phenomena. Simultaneously, a Gaussian mixture background subtraction algorithm is incorporated to supplement the improved three-frame difference method, ensuring that the target image not only has continuous contours but also relatively complete internal information, thereby improving gesture recognition accuracy and reducing false detection rate. Attached Figure Description
[0012] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0013] Figure 1 This is a flowchart of the gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information of the present invention.
[0014] Figure 2 This is a flowchart of the GAN model's workflow. Detailed Implementation
[0015] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0016] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, each technical and scientific term used in these embodiments has the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0017] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0018] In this invention, terms such as "upper," "lower," "left," "right," "front," "back," "vertical," "horizontal," "side," and "bottom" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used only to facilitate the description of the structural relationships of the various components or elements of this invention and do not specifically refer to any component or element in this invention. They should not be construed as limiting the invention.
[0019] In this invention, terms such as "fixed connection," "connected," and "linked" should be interpreted broadly, indicating a fixed connection, an integral connection, or a detachable connection; a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can determine the specific meaning of these terms in this invention based on the specific circumstances, and they should not be construed as limitations on the invention.
[0020] Example 1: like Figure 1 As shown, this embodiment provides a gesture-based haptic interaction method based on the fusion of visual, inertial haptic, and multi-touch information, including the following steps: S1: Acquire video image sequence data of the hand.
[0021] S2: Process the acquired video image sequence data to obtain the target image.
[0022] Specifically, it includes the following steps: S2-1: Obtain the (K-1), K, and K+1th frames of the video image sequence data, and perform grayscale conversion and median filtering on the obtained (K-1), K, and K+1th frames to obtain the video images respectively. , , ; S2-2: Obtain the foreground target image by performing a background difference mixture Gaussian model on the (K+1)th frame of the video image. .
[0023] S2-3: Video images and video images Perform differential processing to obtain the differential image. At the same time, for video images and video images Perform differential processing to obtain the differential image. ; S2-4: For the difference image Image obtained by performing Canny edge detection Simultaneously, for the difference image and Perform a logical AND operation to obtain the image. ; S2-5: Transfer the image and images Perform a logical OR operation to obtain the image. .
[0024] S2-6: Transfer the image and images Perform a logical AND operation to obtain a target image res with a clearer outline and internal information.
[0025] The dynamic gesture recognition method in this embodiment incorporates Canny edge detection into the traditional three-frame difference algorithm and performs a logical OR operation with the three-frame difference images. This not only ensures that the resulting image has good contour information but also helps to compensate for some "holes". Simultaneously, the integration of a Gaussian mixture background subtraction algorithm results in a final target image with continuous contours and relatively complete internal information, improving the recognition rate and reducing the error rate.
[0026] S3: The target image is processed by a generative adversarial network (GAN) model to perform gesture recognition and classification. The recognition and classification results are sent to the human-computer interaction module. The human-computer interaction interface displays the received video image sequence data and displays text and synthesizes speech based on the received recognition results.
[0027] like Figure 2 As shown, the GAN model includes an encoder, decoder, discriminator, and classifier.
[0028] The encoder consists of four 3×3 convolutional layers with a stride of 2, two InceptionV2 layers, and one fully connected layer. CBN layers and Mish activation functions are added to each convolutional layer to normalize the output and enhance the network's expressive power. The fully connected layer adjusts the data dimension to 1×1×100 and outputs the mean and variance to synthesize the sampled variables.
[0029] The decoder consists of one fully connected layer, four transposed convolutional networks, and two InceptionV2-trans networks. Except for the last transposed convolutional layer which uses the Tanh activation function, the rest use the Mish activation function. CBN processing is performed before using the activation function to prevent overfitting.
[0030] The discriminator consists of four convolutional layers and two fully connected layers. Here, convolutional layers with a stride of 2 replace the max pooling layers of the original GAN network to reduce information loss. The Mish activation function is used for the convolutional layers, and the Sigmoid function is used for the fully connected layers. Except for the input layer, all other convolutional layers use CBN for conditional batch normalization.
[0031] The classifier network is isomorphic to the discriminator and is used to determine the image category, but the output of the final fully connected layer is changed from 1 bit to 10 bits, where 10 is the number of gestures to be recognized.
[0032] The gesture recognition of the target image obtained by the generative adversarial network (GAN) model specifically includes the following steps: S3-1: The encoder transmits the input target image (real data). Compressed mapping to the latent space extracts the original features of the target image. ; S3-2: The decoder concatenates the original features z of the target image with the one-hot encoding of the label as input, and outputs the decoded image after decoding. ; S3-3: Transfer the target image and decoded image The input is processed by the discriminator, and the result is output by the Sigmoid layer as an arbitrary probability value between 0 and 1.
[0033] S3-4: Classifier distinguishes target images The category is represented by the probability of each category. The corresponding category can be determined by taking the index of the maximum value of the Argmax function's output.
[0034] Example 2: This embodiment provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the power equipment where the computer-readable storage medium is located to execute the gesture-based somatosensory interaction method described in Embodiment 1, which integrates visual, inertial tactile, and multi-touch information.
[0035] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the terminal embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0036] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can... This can be achieved through other means. For example, the system embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0037] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0038] Additionally, it should be noted that the flowcharts in the accompanying drawings illustrate methods according to embodiments of this disclosure. In the descriptions corresponding to the flowcharts or block diagrams in the drawings, the operations or steps corresponding to different blocks may occur in a different order than disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or sometimes in reverse order, depending on the function involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0039] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information, characterized in that, Includes the following steps: Step 1: Acquire video image sequence data of the hand; Step 2: Process the acquired video image sequence data to obtain the target image; Step 3: Perform gesture recognition and classification on the obtained target image using a generative adversarial network model, and send the recognition and classification results to the human-computer interaction module. The human-computer interaction interface displays the received video image sequence data and displays text and synthesizes speech based on the received recognition results.
2. The gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information as described in claim 1, characterized in that, Step two includes: Acquire the (K-1), K, and K+1th frames of the video image sequence data, and perform grayscale conversion and median filtering on the acquired (K-1), K, and K+1th frames to obtain the video images respectively. , , ; The foreground target image is obtained by performing a background difference mixture Gaussian model on the (K+1)th frame of the video image. ; For video images and video images Perform differential processing to obtain the differential image. At the same time, for video images and video images Perform differential processing to obtain the differential image. ; For difference images Image obtained by performing Canny edge detection Simultaneously, for the difference image and Perform a logical AND operation to obtain the image. ; Image and images Perform a logical OR operation to obtain the image. ; Image and images Perform a logical AND operation to obtain a target image res with a clearer outline and internal information.
3. The gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information as described in claim 1, characterized in that, The generative adversarial network model includes an encoder, a decoder, a discriminator, and a classifier; The encoder is used to process the input target image. Compressed mapping to the latent space extracts the original features of the target image. ; The decoder takes the original features z of the target image and the one-hot encoding of the label as input, decodes them, and outputs the decoded image. ; The discriminator is used to classify the target image. and decoded image The input is processed by the discriminator, and the result is output by the Sigmoid layer as an arbitrary probability value between 0 and 1. The classifier distinguishes the target image. The category is represented by the probability of each category. The corresponding category can be determined by taking the index of the maximum value of the Argmax function's output.
4. The gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information as described in claim 3, characterized in that, The encoder includes four 3×3 convolutional layers with a stride of 2, two InceptionV2 structures, and one fully connected layer. The convolutional layers are equipped with CBN layers and Mish activation functions to normalize the output and enhance the network's expressive power.
5. The gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information as described in claim 3, characterized in that, The decoder includes a fully connected layer, four transposed convolutional networks, and two InceptionV2-trans networks. Among the four transposed convolutional networks, three layers use the Mish activation function, one layer uses the Tanh activation function, and CBN processing is performed before using the activation functions to prevent overfitting.
6. The gesture-based somatosensory interaction method based on the fusion of visual, inertial haptic, and multi-touch information as described in claim 3, is characterized in that... The discriminator consists of four convolutional layers and two fully connected layers. The convolutional layers use the Mish activation function, and the fully connected layers use the Sigmoid function.
7. A computer-readable storage medium comprising a stored program, characterized in that, When the program is running, it controls the power equipment containing the computer-readable storage medium to perform the gesture-based somatosensory interaction method according to any one of claims 1 to 6, which integrates visual, inertial tactile, and multi-touch information.