Human-computer interaction system and method for display equipment
Through deep learning and image processing of hand motion videos, the semantic features of gesture mode images are integrated, and the gesture action intentions are intelligently judged, which solves the problem of unintuitiveness of traditional human-computer interaction systems, and realizes natural and intuitive interaction methods and fast responses.
Patent Information
- Application Number
- CN202411182696.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-08-27
AI Technical Summary
Traditional human-computer interaction systems rely on a fixed-format user interface, are not intuitive or friendly enough, and are difficult to meet the needs of all users. The traditional interaction methods limit the user's range of activities and interaction freedom.
The user's hand movement video is captured through the camera, and the gesture mode image is obtained. The video keyframe analysis technology and image processing algorithm based on deep learning are used to perform feature analysis of hand movement video keyframes, integrate the semantic features of hand movement and gesture mode images, intelligently judge the intention of gesture movement, and return the corresponding interactive content.
It provides a natural and intuitive interaction method, reducing dependence on traditional controllers or input devices, and the system can capture and analyze hand movements in real time, quickly respond to user gestures, and provide a smooth interactive experience.
Smart Images

Figure CN119065502B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent interaction, and more specifically, to a human-computer interaction system and method for displaying equipment. Background Art
[0002] A human-computer interaction system refers to a comprehensive set of technical solutions that allow users to communicate with and control computer systems, machines, or devices through various input methods. These systems are designed to enhance the user experience, making it more intuitive, efficient, and natural. Rapid advances in computer vision, sensor technology, artificial intelligence, and machine learning are making human-computer interaction systems more intelligent and efficient. At the same time, users are increasingly demanding a more natural and intuitive way to interact with their devices.
[0003] However, traditional human-computer interaction systems may rely on fixed user interfaces that are not intuitive or user-friendly, making it difficult to meet the needs of all users. Secondly, traditional interaction methods, such as keystrokes, touch, and mouse clicks, may limit the user's range of activities and freedom of interaction.
[0004] Therefore, an optimized human-computer interaction system for displaying equipment is expected. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a human-computer interaction system and method for displaying equipment, which captures the user's hand movement video through a camera, and obtains a pre-defined gesture pattern image, and adopts a video key frame analysis technology and image processing algorithm based on deep learning to perform feature analysis on each hand movement video key frame, extracts image semantic features from the gesture pattern image, and fuses the feature information of multiple hand movement key frame features in the full time domain with the gesture pattern image semantics, so as to intelligently judge the intention of the gesture action and return the interactive content that matches the intention label of the gesture action. In this way, using hand movements as an interaction means provides users with a natural and intuitive interaction method, reducing dependence on traditional controllers or input devices. At the same time, the system can capture and analyze hand movements in real time to quickly respond to user gestures, thereby providing a smooth interactive experience.
[0006] According to one aspect of the present application, a human-computer interaction system for displaying a device is provided, comprising:
[0007] A hand motion video capture module is used to capture the user's hand motion video through a camera;
[0008] A predefined gesture pattern image acquisition module is used to acquire a predefined gesture pattern image;
[0009] A hand movement key frame image extraction module is used to extract key frames from the user's hand movement video to obtain a time sequence of hand movement key frame images;
[0010] A hand movement key frame temporal feature extraction module, configured to input the time sequence of the hand movement key frame images into a hand movement key feature extractor to obtain a time sequence of a hand movement key feature graph;
[0011] A hand movement key semantic association module is used to extract action key frame semantic features from the time series of the hand movement key feature graph to obtain a time series of hand movement key semantic association feature vectors, and then fuse the time series of the hand movement key semantic association feature vectors to obtain a hand movement key semantic full-time domain association feature vector;
[0012] A predefined gesture pattern semantic extraction module, configured to input the predefined gesture pattern image into a gesture pattern semantic capturer to obtain a predefined gesture pattern semantic feature vector;
[0013] A hand action-gesture pattern semantic feature fusion module, configured to fuse the hand action key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector to obtain a hand action-gesture pattern semantic interaction feature vector;
[0014] The intention judgment result generation module is used to obtain the intention judgment result based on the hand action-gesture pattern semantic interaction feature vector, and return the corresponding interaction content based on the intention judgment result.
[0015] According to another aspect of the present application, a human-computer interaction method for displaying a device is provided, comprising:
[0016] Capturing the user's hand movement video through a camera;
[0017] Get a pre-defined gesture pattern image;
[0018] Extracting key frames from the user's hand movement video to obtain a time sequence of hand movement key frame images;
[0019] Inputting the time series of the hand movement key frame images into a hand movement key feature extractor to obtain a time series of hand movement key feature maps;
[0020] After extracting action key frame semantic features from the time series of the hand action key feature graph to obtain a time series of hand action key semantic association feature vectors, the time series of the hand action key semantic association feature vectors are fused to obtain a hand action key semantic full-time domain association feature vector;
[0021] Inputting the predefined gesture pattern image into a gesture pattern semantics capturer to obtain a predefined gesture pattern semantic feature vector;
[0022] Fusing the hand action key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector to obtain a hand action-gesture pattern semantic interaction feature vector;
[0023] Based on the hand action-gesture pattern semantic interaction feature vector, an intention judgment result is obtained, and based on the intention judgment result, corresponding interaction content is returned.
[0024] Compared with the existing technology, the human-computer interaction system and method for displaying equipment provided by this application captures the user's hand movement video through a camera, obtains a pre-defined gesture pattern image, and uses video key frame analysis technology and image processing algorithms based on deep learning to perform feature analysis on each hand movement video key frame, extracts image semantic features from the gesture pattern image, and fuses the feature information of multiple hand movement key frame features in the full time domain with the gesture pattern image semantics, thereby intelligently judging the intention of the gesture action and returning interactive content that matches the intention label of the gesture action. In this way, using hand movements as an interaction means provides users with a natural and intuitive way of interaction, reducing dependence on traditional controllers or input devices. At the same time, the system can capture and analyze hand movements in real time to quickly respond to user gestures, thereby providing a smooth interactive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0026] Figure 1 4 is a block diagram of a human-computer interaction system for displaying a device according to an embodiment of the present application.
[0027] Figure 2 This is a block diagram of a key semantic association module for hand movements in a human-computer interaction system for displaying a device according to an embodiment of the present application.
[0028] Figure 3 This is a block diagram of an intention judgment result generation module in a human-computer interaction system for displaying a device according to an embodiment of the present application.
[0029] Figure 4Flowchart of a human-computer interaction method for displaying a device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. While the drawings illustrate certain embodiments of the present disclosure, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0031] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0032] It should be noted that the terms "first, second, and third" in the embodiments of the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the terms "first, second, and third" can be interchanged to represent a specific order or precedence where permitted. It should be understood that the objects distinguished by "first, second, and third" can be interchanged where appropriate, such that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0033] As used in this disclosure, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0034] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0035] A human-computer interaction system refers to a comprehensive set of technical solutions that allow users to communicate with and control computer systems, machines, or devices through various input methods. These systems are designed to enhance the user experience, making it more intuitive, efficient, and natural. Rapid advances in computer vision, sensor technology, artificial intelligence, and machine learning are making human-computer interaction systems more intelligent and efficient. At the same time, users are increasingly demanding a more natural and intuitive way to interact with their devices.
[0036] However, traditional human-computer interaction systems may rely on fixed user interfaces that are not intuitive or user-friendly, making it difficult to meet the needs of all users. Secondly, traditional interaction methods, such as keystrokes, touch, and mouse clicks, may limit the user's range of activities and freedom of interaction.
[0037] Therefore, a human-computer interaction system for display equipment is desired, which captures the user's hand movement video through a camera and obtains a pre-defined gesture pattern image, and uses a video key frame analysis technology and image processing algorithm based on deep learning to perform feature analysis on each hand movement video key frame, extracts image semantic features from the gesture pattern image, and fuses the feature information of multiple hand movement key frame features in the full time domain with the gesture pattern image semantics, so as to intelligently judge the intention of the gesture action and return the interactive content that matches the intention label of the gesture action. In this way, using hand movements as an interaction means provides users with a natural and intuitive way of interaction, reducing dependence on traditional controllers or input devices. At the same time, the system can capture and analyze hand movements in real time to quickly respond to user gestures, thereby providing a smooth interactive experience.
[0038] Figure 1 FIG is a block diagram of a human-computer interaction system for displaying a device according to an embodiment of the present application. Figure 1As shown, according to the embodiment of the present application, the human-computer interaction system 100 for displaying equipment includes: a hand motion video capture module 110, which is used to capture the user's hand motion video through a camera; a pre-defined gesture pattern image acquisition module 120, which is used to obtain a pre-defined gesture pattern image; a hand motion key frame image extraction module 130, which is used to perform key frame extraction on the user's hand motion video to obtain a time series of hand motion key frame images; a hand motion key frame temporal feature extraction module 140, which is used to input the time series of the hand motion key frame images into a hand motion key feature extractor to obtain a time series of hand motion key feature maps; a hand motion key semantic association module 150, which is used to perform action key frame semantic feature extraction on the time series of the hand motion key feature maps to obtain hand motion key feature maps. After the time series of the action key semantic association feature vector, the time series of the hand action key semantic association feature vector is fused to obtain the hand action key semantic full-time domain association feature vector; a pre-defined gesture pattern semantic extraction module 160 is used to input the pre-defined gesture pattern image into the gesture pattern semantic capture to obtain the pre-defined gesture pattern semantic feature vector; a hand action-gesture pattern semantic feature fusion module 170 is used to fuse the hand action key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector to obtain the hand action-gesture pattern semantic interaction feature vector; an intention judgment result generation module 180 is used to obtain the intention judgment result based on the hand action-gesture pattern semantic interaction feature vector, and return the corresponding interaction content based on the intention judgment result.
[0039] In an embodiment of the present application, the hand motion video capture module 110 is used to capture the user's hand motion video through a camera. It should be understood that the user's hand motion video refers to a continuous image sequence of the user's hand movement and action in space recorded by a camera. Based on this, in order to more clearly capture the user's interaction intention and provide more accurate interaction results, in the technical solution of the present application, the user's hand motion video is captured by a camera, and the hand motion video is subjected to feature analysis and processing, which can help the system intelligently determine the operation or interaction the user wants to perform, thereby providing the user with a smooth interaction experience.
[0040] In an embodiment of the present application, the pre-defined gesture pattern image acquisition module 120 is used to acquire a pre-defined gesture pattern image. Accordingly, it is considered that the pre-defined gesture pattern image refers to a standard image or model of a gesture that is pre-set and stored in the human-computer interaction system. These images represent specific gesture actions that the system can recognize and respond to. Therefore, in the technical solution of the present application, the pre-defined gesture pattern image is obtained, and the gesture pattern image can be used as a known standard or template to compare with the user's actual hand movements, thereby helping the system to recognize and verify specific gestures.
[0041] In an embodiment of the present application, the hand movement key frame image extraction module 130 is used to perform key frame extraction on the hand movement video of the user to obtain a time sequence of hand movement key frame images. It should be understood that, considering that the hand movement video of the user usually contains the most critical and important hand movement interaction feature information of the user's hand movement. Therefore, in order to be able to extract key hand movement image frame information from the hand movement video of the user, in the technical solution of the present application, the hand movement video of the user is subjected to key frame extraction to obtain a time sequence of hand movement key frame images. That is, key frames usually contain the most important information about the user's actions or events when performing gesture interaction, which helps to quickly identify and understand gestures, and key frames are the most representative frames in the video, which can reduce the amount of data that needs to be processed, thereby simplifying the analysis process and obtaining a time sequence of hand movement key frame images with the most feature expression capabilities.
[0042] In an embodiment of the present application, the hand movement key frame temporal feature extraction module 140 is used to input the time series of the hand movement key frame images into the hand movement key feature extractor to obtain the time series of the hand movement key feature graph. Specifically, in an embodiment of the present application, the hand movement key frame temporal feature extraction module is used to: input the time series of the hand movement key frame images into the hand movement key feature extractor based on the residual neural network model to obtain the time series of the hand movement key feature graph. In particular, the hand movement key feature extractor based on the residual neural network model here is essentially a residual neural network model. Accordingly, considering that each hand movement key frame image in the time series of the hand movement key frame images contains the key movement feature information of the hand movement, and the residual neural network model can learn feature representations from low-level to high-level, this is very important for understanding the key details and context of the hand movement. Based on this, in the technical solution of the present application, the time series of the hand movement key frame images is input into the hand movement key feature extractor based on the residual neural network model to obtain the time series of the hand movement key feature graph. That is, the residual network reuses the features of the previous layers through skip connections, which helps to maintain the important initial features of hand movements in the deep network, thereby extracting richer and more abstract feature representations.
[0043] In an embodiment of the present application, the hand movement key semantic association module 150 is used to perform action key frame semantic feature extraction on the time series of the hand movement key feature graph to obtain a time series of hand movement key semantic association feature vectors, and then fuse the time series of the hand movement key semantic association feature vectors to obtain a hand movement key semantic full-time domain association feature vector. Figure 2 : is a block diagram of a key semantic association module for hand movements in a human-computer interaction system for displaying a device according to an embodiment of the present application. Specifically, in the embodiment of the present application, Figure 2 As shown, the hand movement key semantic association module 150 includes: a hand movement key frame image semantic feature extraction unit 151, which is used to input the time series of the hand movement key feature map into the VIT model to perform key frame semantic feature extraction to obtain the time series of the hand movement key semantic association feature vector; a hand movement key frame image full time domain semantic association unit 152, which is used to cascade the time series of the hand movement key semantic association feature vector to obtain the hand movement key semantic full time domain association feature vector.
[0044] Specifically, the hand movement key frame image semantic feature extraction unit 151 is used to input the time series of the hand movement key feature map into the VIT model for key frame semantic feature extraction to obtain the time series of the hand movement key semantic association feature vector. It should be understood that, considering the existence of contextual semantic association relationships about local features between different channel features in the hand movement key feature map, and the VIT model adopts the attention mechanism in the Transformer architecture, this enables it to effectively capture the correlation between different areas in the image, thereby better understanding the semantic information of the hand movement. Therefore, in the technical solution of the present application, the time series of the hand movement key feature map is input into the VIT model for key frame semantic feature extraction to capture and extract the semantic association relationships between the contexts between each channel, thereby obtaining the time series of the hand movement key semantic association feature vector.
[0045] Specifically, the hand action key frame image full-time domain semantic association unit 152 is used to cascade the time series of the hand action key semantic association feature vectors to obtain the hand action key semantic full-time domain association feature vectors. Furthermore, the time series of the hand action key semantic association feature vectors are cascaded to obtain the hand action key semantic full-time domain association feature vectors. This can maintain the temporal continuity of the action features, which is crucial for understanding the dynamic changes and development trends of the action, and can integrate the hand action semantic information at different time points to provide a comprehensive feature representation, which helps to capture comprehensive key information in the entire action process.
[0046] In an embodiment of the present application, the pre-defined gesture pattern semantic extraction module 160 is used to input the pre-defined gesture pattern image into a gesture pattern semantic capturer to obtain a pre-defined gesture pattern semantic feature vector. Specifically, in an embodiment of the present application, the pre-defined gesture pattern semantic extraction module is used to: input the pre-defined gesture pattern image into a gesture pattern semantic capturer based on a densely connected network to obtain the pre-defined gesture pattern semantic feature vector. Accordingly, considering that the pre-defined gesture pattern image expresses the defined gesture pattern information, and the pattern information contains key gesture image feature information, and considering that the densely connected network improves the information flow and feature reuse of the network by connecting each layer to all previous layers, it helps to extract richer features. Based on this, in the technical solution of the present application, the pre-defined gesture pattern image is input into a gesture pattern semantic capturer based on a densely connected network to obtain the pre-defined gesture pattern semantic feature vector. The gesture pattern semantic capturer based on a densely connected network here is essentially a densely connected neural network model. In particular, compared with traditional convolutional networks, densely connected networks reduce the loss of information during network transmission, help retain more original features of gesture patterns, and thus extract accurate semantic features, providing strong support for gesture recognition.
[0047] In an embodiment of the present application, the hand movement-gesture pattern semantic feature fusion module 170 is used to fuse the hand movement key semantic full-time domain associated feature vector and the pre-defined gesture pattern semantic feature vector to obtain a hand movement-gesture pattern semantic interaction feature vector. Specifically, in an embodiment of the present application, the hand movement-gesture pattern semantic feature fusion module includes: a pre-defined gesture pattern semantic feature distribution deviation quantization unit, which is used to perform feature value expression enhancement based on feature distribution deviation quantization on the pre-defined gesture pattern semantic feature vector to obtain an optimized pre-defined gesture pattern semantic feature vector; a hand movement-gesture pattern semantic fusion unit, which is used to fuse the hand movement key semantic full-time domain associated feature vector and the optimized pre-defined gesture pattern semantic feature vector to obtain the hand movement-gesture pattern semantic interaction feature vector.
[0048] Specifically, the pre-defined gesture pattern semantic feature distribution deviation quantization unit is used to enhance the feature value expression of the pre-defined gesture pattern semantic feature vector based on the feature distribution deviation quantization to obtain an optimized pre-defined gesture pattern semantic feature vector. In particular, in the technical solution of the present application, it is considered that the hand movement key frame image is extracted from the user's real-time hand movement video, while the gesture pattern image is a pre-defined static image. The collection methods and scenarios of these two types of data may be different, resulting in differences in feature distribution. The hand movement key feature map is extracted through a residual neural network model, and then semantic features are extracted through a VIT model; while the gesture pattern image extracts semantic features through a densely connected network capturer. Different model architectures may lead to differences in feature representation. The hand movement key frame image has time series characteristics, reflecting the dynamic changes of the movement; while the gesture pattern image is static and does not contain time series information. This difference between time series and static features may lead to inconsistent feature distribution. However, considering that the pre-defined gesture pattern semantic feature vector has a feature distribution shift deviation relative to the hand action key semantic full-time domain association feature vector, this causes the fused hand-gesture pattern semantic interaction feature vector to have an extreme category offset when it is classified and judged by the classifier-based hand interaction intention judge, resulting in a decrease in the accuracy of the classification result. Based on this, in the technical solution of the present application, the pre-defined gesture pattern semantic feature vector is enhanced by the feature value expression based on the quantization of the feature distribution deviation to obtain an optimized pre-defined gesture pattern semantic feature vector.
[0049] More specifically, the pre-defined gesture pattern semantic feature distribution deviation quantization unit is used to: calculate the shift difference representation vector between the hand movement key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector; perform probabilistic activation function on each eigenvalue of the shift difference representation vector to obtain a probabilistic shift difference representation vector; input the pre-defined gesture pattern semantic feature vector into the classifier-based hand interaction intention judge to obtain an initial classification probability value; determine the backpropagation compensation representation vector based on the comparison between the eigenvalues of each position of the probabilistic shift difference representation vector and the initial classification probability value and the weight matrix of the classifier-based hand interaction intention judge; calculate the exponential function value of each eigenvalue of the backpropagation compensation representation vector to obtain a backpropagation compensation class representation vector; calculate the positional point multiplication between the backpropagation compensation class representation vector and the pre-defined gesture pattern semantic feature vector to obtain a point product result vector, and calculate the weighted sum between the point product result vector and the pre-defined gesture pattern semantic feature vector to obtain the optimized pre-defined gesture pattern semantic feature vector.
[0050] In the embodiment of the present application, specifically, the feature value expression enhancement based on the feature distribution deviation quantification of the pre-defined gesture pattern semantic feature vector is performed using the following feature value expression enhancement formula to obtain an optimized pre-defined gesture pattern semantic feature vector; wherein, the feature value expression enhancement formula is:
[0051]
[0052] Among them, Sigmoid represents the sigmoid function, X a represents the full-time domain correlation feature vector of the key semantics of the hand action, X b represents the semantic feature vector of the predefined gesture pattern, X d represents the probabilistic shift difference representation vector, represents the eigenvalue of the i-th position in the probabilistic shifted differential representation vector, sgn represents the sign function, δ and γ represent the first and second weighted hyperparameters respectively, and M c represents the weight feature matrix of the classifier-based hand interaction intention judgement, θ represents the initial classification probability value, ⊙ represents the position point multiplication, represents matrix multiplication, X b ' represents the optimized pre-defined gesture pattern semantic feature vector.
[0053] That is, to address the distribution shift deviation problem between feature vectors, the pre-defined gesture pattern semantic feature vector is enhanced based on the feature distribution deviation quantification based on the shift information between the hand action key semantic full-time domain associated feature vector and the pre-defined gesture pattern semantic feature vector. The shift difference representation vector is obtained by calculating the difference between the hand action key semantic full-time domain associated feature vector and the pre-defined gesture pattern semantic feature vector to quantify the feature distribution deviation. The shift difference representation vector reveals the alignment of the two feature vectors in the feature space, providing a basis for subsequent compensation. Then, a probability activation function is applied to each eigenvalue of the shift difference representation vector for probabilistic processing, thereby converting the difference value into a probability distribution, allowing the model to understand the deviation between features in a probabilistic form. The original pre-defined gesture pattern semantic feature vector is then input into the classifier-based hand interaction intention determiner to obtain an initial classification probability value, wherein the initial classification probability value provides the model with a basic classification performance indicator before feature compensation. Next, the backpropagation compensation representation vector is determined by combining the probabilistic shifted difference representation vector with the initial classification probability and the weight matrix of the classifier-based hand interaction intention predictor. Specifically, the model's own weights and bias information are used to compensate the feature vector to reduce classification error. Next, an exponential function is calculated for each eigenvalue of the backpropagation compensation representation vector, enhancing the range of the eigenvalues and providing a mechanism for scaling feature compensation. Finally, the optimized pre-defined gesture pattern semantic feature vector is obtained by calculating the dot product between the backpropagation compensation class representation vector and the pre-defined gesture pattern semantic feature vector, and then calculating a weighted sum with the original pre-defined gesture pattern semantic feature vector. It should be understood that the dot product and weighted sum operations combine the compensation information with the original feature information to generate an adjusted feature vector, aiming to reduce extreme class deviations and improve classification accuracy. The entire data processing process achieves the goal of identifying feature distribution shifts, compensating for these deviations, and ultimately optimizing the feature representation. This process not only improves the adaptability of the classifier-based hand interaction intention predictor to changes in feature distribution but also enhances the model's ability to distinguish between different classes. Through the feature value expression enhancement mechanism based on the quantification of feature distribution deviation, the model can more accurately capture the subtle differences between feature vectors, thereby achieving higher accuracy and better generalization performance in classification tasks.
[0054] Specifically, the hand movement-gesture pattern semantic fusion unit is used to fuse the hand movement key semantic full-time domain associated feature vector and the optimized pre-defined gesture pattern semantic feature vector to obtain the hand movement-gesture pattern semantic interaction feature vector. It should be understood that the hand movement key semantic full-time domain associated feature vector captures the overall semantic information of the user's hand movement changing over time. The pre-defined gesture pattern semantic feature vector contains the visual and semantic information of the gesture. Therefore, in order to provide a more comprehensive description of hand movements and increase the accuracy of gesture recognition, in the technical solution of the present application, the hand movement key semantic full-time domain associated feature vector and the pre-defined gesture pattern semantic feature vector are fused to obtain the hand movement-gesture pattern semantic interaction feature vector, thereby realizing a more advanced human-computer interaction function.
[0055] In an embodiment of the present application, the intention judgment result generation module 180 is used to obtain an intention judgment result based on the hand action-gesture pattern semantic interaction feature vector, and return corresponding interaction content based on the intention judgment result. Figure 3 : is a block diagram of an intention judgment result generation module in a human-computer interaction system for displaying a device according to an embodiment of the present application. Specifically, in the embodiment of the present application, Figure 3 As shown, the intention judgment result generation module 180 includes: an intention label judgment unit 181, which is used to input the hand action-gesture pattern semantic interaction feature vector into a classifier-based hand interaction intention judgement to obtain the intention label of the gesture action; an interaction content return unit 182, which is used to return the interaction content that matches the intention label of the gesture action based on the intention label of the gesture action. That is, the hand action-gesture pattern semantic interaction feature vector is used for classification processing, so as to intelligently judge the intention of the gesture action and return the interaction content that matches the intention label of the gesture action. In this way, using hand actions as an interaction means provides users with a natural and intuitive interaction method, reducing dependence on traditional controllers or input devices. At the same time, the system can capture and analyze hand actions in real time to quickly respond to user gestures, thereby providing a smooth interaction experience.
[0056] In summary, a human-computer interaction system 100 for displaying a device based on an embodiment of the present application is explained, which captures a user's hand movement video through a camera, obtains a pre-defined gesture pattern image, and uses a video key frame analysis technology and image processing algorithm based on deep learning to perform feature analysis on each hand movement video key frame, extracts image semantic features from the gesture pattern image, and fuses the feature information of multiple hand movement key frame features in the full time domain with the gesture pattern image semantics, thereby intelligently judging the intention of the gesture action and returning interactive content that matches the intention label of the gesture action. In this way, using hand movements as an interaction means provides users with a natural and intuitive way of interaction, reducing dependence on traditional controllers or input devices. At the same time, the system can capture and analyze hand movements in real time to quickly respond to user gestures, thereby providing a smooth interactive experience.
[0057] Figure 4 Flowchart of the human-computer interaction method for displaying a device according to an embodiment of the present application. Figure 4 As shown, the human-computer interaction method for displaying a device according to an embodiment of the present application includes: S110, capturing a user's hand movement video through a camera; S120, obtaining a pre-defined gesture pattern image; S130, performing key frame extraction on the user's hand movement video to obtain a time sequence of hand movement key frame images; S140, inputting the time sequence of the hand movement key frame images into a hand movement key feature extractor to obtain a time sequence of hand movement key feature maps; S150, performing action key frame semantic feature extraction on the time sequence of the hand movement key feature maps to obtain a time sequence of hand movement key semantic association feature vectors. After the time sequence, the time series of the hand movement key semantic association feature vector is fused to obtain the hand movement key semantic full-time domain association feature vector; S160, the pre-defined gesture pattern image is input into the gesture pattern semantic capturer to obtain the pre-defined gesture pattern semantic feature vector; S170, the hand movement key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector are fused to obtain the hand movement-gesture pattern semantic interaction feature vector; S180, based on the hand movement-gesture pattern semantic interaction feature vector, the intention judgment result is obtained, and based on the intention judgment result, the corresponding interaction content is returned.
[0058] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned human-computer interaction method for displaying a device have been described in detail in the above reference. Figures 1 to 3 The human-computer interaction system for displaying the device has been described in detail, and therefore, its repeated description will be omitted.
[0059] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0060] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0061] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0062] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in hardware, or, as described above, those skilled in the art will readily appreciate that the present invention may be implemented in hardware, firmware, or a combination thereof. When implemented using software, the aforementioned functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media includes any medium that facilitates the transfer of computer programs from one location to another. Storage media may be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection may appropriately constitute a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of protection of computer-readable media.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit of the technical solutions of the present invention.
Claims
1. A human-computer interaction system for displaying equipment, characterized in that: include: A hand motion video capture module is used to capture the user's hand motion video through a camera; A predefined gesture pattern image acquisition module is used to acquire a predefined gesture pattern image; A hand movement key frame image extraction module is used to extract key frames from the user's hand movement video to obtain a time sequence of hand movement key frame images; A hand movement key frame temporal feature extraction module, configured to input the time sequence of the hand movement key frame images into a hand movement key feature extractor to obtain a time sequence of a hand movement key feature graph; A hand movement key semantic association module is used to extract action key frame semantic features from the time series of the hand movement key feature graph to obtain a time series of hand movement key semantic association feature vectors, and then fuse the time series of the hand movement key semantic association feature vectors to obtain a hand movement key semantic full-time domain association feature vector; A predefined gesture pattern semantic extraction module, configured to input the predefined gesture pattern image into a gesture pattern semantic capturer to obtain a predefined gesture pattern semantic feature vector; A hand action-gesture pattern semantic feature fusion module, configured to fuse the hand action key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector to obtain a hand action-gesture pattern semantic interaction feature vector; an intention judgment result generating module, configured to obtain an intention judgment result based on the hand action-gesture pattern semantic interaction feature vector, and return corresponding interaction content based on the intention judgment result; The hand movement-gesture pattern semantic feature fusion module includes: a predefined gesture pattern semantic feature distribution deviation quantization unit, configured to perform feature value expression enhancement based on feature distribution deviation quantization on the predefined gesture pattern semantic feature vector to obtain an optimized predefined gesture pattern semantic feature vector; A hand action-gesture pattern semantic fusion unit, configured to fuse the hand action key semantic full-time domain association feature vector and the optimized pre-defined gesture pattern semantic feature vector to obtain the hand action-gesture pattern semantic interaction feature vector; The gesture pattern semantic feature distribution deviation quantification unit is defined in advance and is used to: Calculating a shift difference representation vector between the hand action key semantic full-time domain associated feature vector and the pre-defined gesture pattern semantic feature vector; Probabilistically converting each eigenvalue of the shifted differential representation vector based on a probability activation function to obtain a probabilistic shifted differential representation vector; Inputting the predefined gesture pattern semantic feature vector into a hand interaction intention determiner based on a classifier to obtain an initial classification probability value; Determining a back-propagation compensation representation vector based on a comparison between the eigenvalues of each position of the probabilistic shifted differential representation vector and the initial classification probability value and a weight matrix of the classifier-based hand interaction intention determiner; Calculating the exponential function value of each eigenvalue of the backpropagation compensation representation vector to obtain a backpropagation compensation class representation vector; Calculate the positional point multiplication between the backpropagation compensation class representation vector and the pre-defined gesture pattern semantic feature vector to obtain a point product result vector, and calculate the weighted sum between the point product result vector and the pre-defined gesture pattern semantic feature vector to obtain the optimized pre-defined gesture pattern semantic feature vector.
2. The human-computer interaction system for displaying equipment according to claim 1, characterized in that: The hand action key frame temporal feature extraction module is used to: The time series of the hand movement key frame images is input into a hand movement key feature extractor based on a residual neural network model to obtain the time series of the hand movement key feature graphs.
3. The human-computer interaction system for displaying equipment according to claim 2, characterized in that: The hand movement key semantic association module includes: A hand movement key frame image semantic feature extraction unit is used to input the time series of the hand movement key feature map into the VIT model to perform key frame semantic feature extraction to obtain the time series of the hand movement key semantic association feature vector; The hand movement key frame image full time domain semantic association unit is used to cascade the time series of the hand movement key semantic association feature vector to obtain the hand movement key semantic full time domain association feature vector.
4. The human-computer interaction system for displaying equipment according to claim 3, characterized in that: The predefined gesture pattern semantic extraction module is used to: The pre-defined gesture pattern image is input into a gesture pattern semantic capturer based on a densely connected network to obtain the pre-defined gesture pattern semantic feature vector.
5. The human-computer interaction system for displaying equipment according to claim 4, characterized in that: The intention judgment result generation module includes: an intention label judgment unit, configured to input the hand action-gesture pattern semantic interaction feature vector into the classifier-based hand interaction intention judgement unit to obtain an intention label of the gesture action; The interactive content returning unit is configured to return interactive content matching the intention tag of the gesture action based on the intention tag of the gesture action.
6. A human-computer interaction method for displaying equipment, using the human-computer interaction system for displaying equipment according to claim 1, characterized in that: include: Capturing the user's hand movement video through a camera; Get a pre-defined gesture pattern image; Extracting key frames from the user's hand movement video to obtain a time sequence of hand movement key frame images; Inputting the time series of the hand movement key frame images into a hand movement key feature extractor to obtain a time series of hand movement key feature maps; After extracting action key frame semantic features from the time series of the hand action key feature graph to obtain a time series of hand action key semantic association feature vectors, the time series of the hand action key semantic association feature vectors are fused to obtain a hand action key semantic full-time domain association feature vector; Inputting the predefined gesture pattern image into a gesture pattern semantics capturer to obtain a predefined gesture pattern semantic feature vector; Fusing the hand action key semantic full-time domain association feature vector and the pre-defined gesture pattern semantic feature vector to obtain a hand action-gesture pattern semantic interaction feature vector; Based on the hand action-gesture pattern semantic interaction feature vector, an intention judgment result is obtained, and based on the intention judgment result, corresponding interaction content is returned.
7. The human-computer interaction method for display equipment according to claim 6, characterized in that: Inputting the time sequence of the hand movement key frame images into a hand movement key feature extractor to obtain a time sequence of the hand movement key feature graph, comprising: The time series of the hand movement key frame images is input into a hand movement key feature extractor based on a residual neural network model to obtain the time series of the hand movement key feature graphs.
8. The human-computer interaction method for display equipment according to claim 7, characterized in that: Inputting the predefined gesture pattern image into a gesture pattern semantics capturer to obtain a predefined gesture pattern semantic feature vector, including: The pre-defined gesture pattern image is input into a gesture pattern semantic capturer based on a densely connected network to obtain the pre-defined gesture pattern semantic feature vector.
Citation Information
Patent Citations
Wind power generation equipment immersive simulation practical training system based on virtual reality
CN118535021A